Background: Physicians obtain answers to less than half of the clinical questions they ask at bedside and on rounds (1,2). Generative AI holds the potential to narrow this gap, but few AI tools have been trained to address authentic clinical questions by hospitalists and their teams. We created VLRChat, a custom AI application that draws from a publicly available clinical textbook (Internal Medicine Housestaff Handbook). Through retrieval-augmented generation (RAG), VLRChat provides accurate, relevant, and trustworthy answers to clinical questions asked by clinicians and learners (3,4). During a pilot phase from February to July 2025, VLRChat collected over 1000 questions from medical students, residents, and attendings. Analysis of this dataset of clinical question-answer pairs can provide new insights to clinical care and medical education (5).

Purpose: The purpose of the present study is to systematically analyze clinical questions submitted to VLRChat, with a focus on identifying educational gaps and user-specific needs.

Description: This is a retrospective study of existing, de-identified VLRChat interaction logs from both internal and external users who utilized the experimental tool via our publicly available webpage (Vimbook.org). Data was stored in a password protected database, and no patient identifying information was collected. Questions were randomly sampled and analyzed based on their characteristics, such as taxonomy and specialty. We analyzed the most common words and phrases in questions. Timestamp analysis was also performed.

Conclusions: We randomly sampled 250 questions; 33 questions were irrelevant or resulted from a connection error. Out of the 217 usable questions, 28 questions were non-clinical in nature (e.g. Door code, logistics), and 189 questions were clinical. Out of the clinical questions, the most frequently asked topics were in Cardiology (38/189, 20%), Infectious Diseases (33/189, 17%), Gastroenterology (33/189, 17%), and most common taxonomies involved treatment (93/189, 49%), diagnosis (44/189, 23%), and etiology (19/189, 10%). There were 8 questions on education and 5 on patient communication. The most frequently asked clinical words were hypertension (16), antibiotics (13), catatonia (8), and pancreatitis (5). The most common phrases were “treat hypertension” (14), “social determinants” (4), “acute pancreatitis” (3), “diabetes management” (3). 42 questions were asked after working hours between 7 PM to 7AM, 69 questions were asked in the morning between 7AM to 12PM, and 106 questions were asked in the afternoon between 12PM to 7PM.The data generated by clinical questions in hospital medicine provides valuable insights into medical education and training to ultimately enhance patient care decisions. When questions are not fully answered in real-time during rounds, they are often brought to VLRChat. This has enabled the creation of a record of what people are asking in real-time. Analyzing patterns that repeatedly show up such as specific words, diagnoses, specialties, or topics can help identify educational gaps and system needs, at department or institutional level. Applied AI tools with domain- and institution-specific knowledge such as VLRChat will play a key role in the clinical learning environment of the future.

IMAGE 1: Figure 1. Architecture Components of RAG-Enhanced AI Chatbot, VLRChat

IMAGE 2: Figure 2. Breakdown of Authentic Clinical Questions Asked of VLRChat by Subtype