Can a chatbot help steer a biopsy needle through the liver? A new pilot study from researchers at Memorial Sloan Kettering Cancer Center suggests that, in a carefully controlled and highly artificial setting, the answer is a cautious maybe. The team, led by Francois H. Cornelis with first author Isabelle Echelman, tested whether GPT-4.1, OpenAI’s multimodal generative AI model, could recommend safe entry angles for CT-guided liver needle biopsies without any case-specific training. Their findings, published open access in CVIR Oncology, offer one of the first glimpses of how general-purpose generative AI might handle one of interventional radiology’s most delicate planning tasks, while also laying bare just how far the technology remains from the operating suite.
Liver biopsy is a cornerstone of cancer diagnosis, but it is far from trivial. The liver is a crowded, vascular organ hemmed in by ribs, lungs, bowel, and major blood vessels, and the needle must reach a target lesion while avoiding all of these hazards. Today, that planning rests almost entirely on the experience and judgment of individual interventional radiologists, which introduces meaningful variability between operators. Standardized AI-assisted planning has been proposed as a way to reduce that variability, particularly at centers with less experienced staff, but the field remains preliminary. The Memorial Sloan Kettering team set out to test the most basic building block of such a system: could a generative AI model, given nothing but a single axial CT slice and a standardized prompt, propose an entry angle that experienced reviewers would rate as potentially safe?
The study design was deliberately simple. The researchers retrospectively collected de-identified CT images from 30 consecutive CT-guided liver biopsies performed between February and June 2025 by sixteen board-certified interventional radiologists with a median of eight years of experience. The patient group included 20 women and 10 men with a median age of 56, and their lesions spanned liver segments 1 through 8 under the Couinaud classification. Crucially, the AI never touched a patient. It worked only on the initial planning CT image, after the fact, and only in two dimensions, while the actual procedures had already been completed safely by human hands.
The preprocessing workflow was strikingly low-tech for a study of cutting-edge AI. Each CT image was cropped to remove patient identifiers, magnified without rescaling, and overlaid with a radial grid created in GIMP, the free open-source image editor. The grid consisted of 36 radial lines at 10-degree intervals with concentric circles spaced 10 millimeters apart, calibrated to known anatomical landmarks. The annotated images were saved as PNG files and fed to GPT-4.1 through a standardized prompt. The model operated entirely on its baseline training, with no fine-tuning on medical data, and a fresh chat was started for every case to prevent contamination between cases.
The researchers tested two conditions that differed in how much hand-holding the AI received. In the first condition, two study authors manually annotated each image with a blue circle marking the lesion and a red line delineating sensitive structures to avoid, essentially handing the model a pre-drawn map of the danger zones. In the second, more demanding condition, only the blue circle was drawn, forcing the AI to identify critical structures on its own. For each case, the model was allowed up to four iterations: if its proposed trajectory failed safety criteria, it was re-prompted within that limit. Two independent reviewers then rated every AI-generated trajectory using consensus methodology, judging safety, avoidance of organs, vessels, and bone, whether the path crossed the midline or excessive tissue, and whether the approach would be comfortable and clinically reasonable.
The headline numbers were encouraging, with important caveats. In the first condition, the AI needed a median of two iterations and produced trajectories rated potentially safe in 27 of 30 cases, or 90 percent. In the second condition, without the red-line annotations, it needed a median of just one iteration and scored 28 of 30, or 93 percent. The difference between conditions was not statistically significant, and the model’s proposed entry angles differed significantly from the angles actually used by clinicians, with mean differences reaching statistical significance in both conditions. Only 23 percent of cases in condition one and 33 percent in condition two showed strong concordance, defined as entry angles within 10 degrees with similar liver traversal patterns, meaning the AI was often safe but not necessarily matching human choices.
The failures are as instructive as the successes. In three cases in the first condition, the AI proposed paths crossing bone, the midline, or the marked red line, even after four attempts. One of these involved a patient with ascites, fluid in the abdomen that complicated the anatomy; another involved a patient whose prior surgery had removed the left hepatic lobe, creating atypical anatomy; and a third targeted a high-lying segment 2 lesion where the model routed the needle through costal cartilage. Cartilage proved a recurring blind spot: in the second condition, both unsafe trajectories involved bone or cartilage, suggesting the model struggles to recognize structures with intermediate radiodensity. The authors note that excluding these inadequate cases from distance calculations introduces selection bias that may flatter the AI’s performance.
On the metric of liver tissue traversed, the AI actually looked competitive. In condition one, its paths crossed a median of 31.6 millimeters of liver versus 48.25 millimeters for the manual paths, and in condition two, 39.35 versus 48.25 millimeters, though neither difference reached statistical significance. Interestingly, the AI’s longer median distances in the unannotated condition hint at a conservative streak: without predefined avoidance zones, it may favor longer but safer routes. A power analysis indicated the 30-case sample could only detect large effect sizes, so smaller but clinically meaningful differences would require larger cohorts to detect.
The authors are refreshingly blunt about what these results do not mean. This was a proof of concept, nothing more. The workflow does not address a real clinical need: interventional radiologists can identify safe trajectories in seconds using native console capabilities with full three-dimensional visualization, whereas the AI pipeline required manual image export, GIMP annotation, and chatbot prompting, all while discarding the three-dimensional spatial information that real planning depends on. A human planner scrolls through multiple slices, uses multiplanar reconstruction, and mentally constructs 3D anatomy, capabilities entirely absent from a static 2D analysis. The AI also required continuous human supervision, with iterative re-prompting in most cases, which undermines any claim of autonomy. And because the AI trajectories were never actually executed, they remain theoretical, while the manual paths were proven safe by successful, complication-free procedures, a comparison that cannot establish equivalence.
Looking forward, the researchers see a path, if a long one. Future systems would need automated lesion identification, since current generative models lack the segmentation capabilities required for precise localization, and integration with dedicated segmentation architectures such as U-Net could enable fully autonomous planning from raw CT images. Three-dimensional volumetric analysis, quantitative vascular proximity measurements, phantom studies, and prospective multicenter trials would all be prerequisites before any clinical consideration. The authors also point toward applications in tumor ablation and margin assessment, and note that many liver biopsies worldwide are performed under ultrasound guidance, where real-time, operator-dependent imaging poses an even harder challenge for AI. For now, the message is measured: generative AI can look at a liver CT slice and propose a plausible needle angle most of the time, but the 7 to 10 percent of cases where it fails, sometimes by routing a needle through bone or cartilage, are exactly the cases where a patient would be harmed. In medicine, that margin is everything.
Subject of Research: Feasibility of generative AI for planning CT-guided liver biopsy needle trajectories
Article Title: Feasibility of generative AI for CT-guided liver biopsy trajectory planning: a pilot study
Article References: Feasibility of generative AI for CT-guided liver biopsy trajectory planning: a pilot study. (n.d.). https://doi.org/10.1007/s44343-026-00041-7
Image Credits: AI Generated
DOI: 10.1007/s44343-026-00041-7
Keywords: generative AI, GPT-4.1, liver biopsy, CT guidance, interventional radiology, trajectory planning, pilot study, medical imaging, needle biopsy, proof of concept, Memorial Sloan Kettering, artificial intelligence
Cite Scienmag News
Nathaniel Bowman. (September 24, 2026). AI Chatbot Plans Liver Biopsy Needle Paths in Landmark Pilot Test. Scienmag. https://scienmag.com/ai-chatbot-plans-liver-biopsy-needle-paths-in-landmark-pilot-test/
Nathaniel Bowman. "AI Chatbot Plans Liver Biopsy Needle Paths in Landmark Pilot Test." Scienmag, 24 September 2026, https://scienmag.com/ai-chatbot-plans-liver-biopsy-needle-paths-in-landmark-pilot-test/. Accessed 24 September 2026.
Nathaniel Bowman. "AI Chatbot Plans Liver Biopsy Needle Paths in Landmark Pilot Test." Scienmag. September 24, 2026. https://scienmag.com/ai-chatbot-plans-liver-biopsy-needle-paths-in-landmark-pilot-test/

