<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>method &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/method/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Tue, 22 Sep 2026 14:13:36 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>method &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>Small Models, Big Payoff: Teamwork Fixs AI Tool-Calling Errors</title>
		<link>https://scienmag.com/small-models-big-payoff-teamwork-fixs-ai-tool-calling-errors/</link>
		
		<dc:creator><![CDATA[Violet Maxwell]]></dc:creator>
		<pubDate>Tue, 22 Sep 2026 14:13:36 +0000</pubDate>
				<category><![CDATA[Earth Science]]></category>
		<category><![CDATA[agentic AI systems]]></category>
		<category><![CDATA[AI system bottlenecks]]></category>
		<category><![CDATA[AI task planning]]></category>
		<category><![CDATA[AI tool invocation errors]]></category>
		<category><![CDATA[API request formatting]]></category>
		<category><![CDATA[autonomous AI task execution]]></category>
		<category><![CDATA[collaboration]]></category>
		<category><![CDATA[collaboration between large and small models]]></category>
		<category><![CDATA[Enhanced]]></category>
		<category><![CDATA[improving AI tool accuracy]]></category>
		<category><![CDATA[invocation]]></category>
		<category><![CDATA[large language models]]></category>
		<category><![CDATA[method]]></category>
		<category><![CDATA[multi-model]]></category>
		<category><![CDATA[natural language to machine commands]]></category>
		<category><![CDATA[Scientific Research]]></category>
		<category><![CDATA[small models for AI]]></category>
		<category><![CDATA[tool]]></category>
		<category><![CDATA[tool selection in AI]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=205719</guid>

					<description><![CDATA[Large language models have dazzled the world with their ability to write, reason and converse, but when it comes to actually doing things—booking a flight, querying a database, triggering a smart-home routine—they often stumble on something almost embarrassingly mundane: formatting.]]></description>
										<content:encoded><![CDATA[<p>Large language models have dazzled the world with their ability to write, reason and converse, but when it comes to actually doing things—booking a flight, querying a database, triggering a smart-home routine—they often stumble on something almost embarrassingly mundane: formatting. A new study published in the open-access journal Vicinagearth argues that the single biggest bottleneck in letting AI agents call external tools is not intelligence at all, but the rigid syntactic discipline required to produce a machine-readable API request. The research team, led by Yudian Zhang and Xuelong Li at the Institute of Artificial Intelligence (TeleAI) of China Telecom, together with Haijiang Zhu of Beijing University of Chemical Technology, proposes an elegantly simple remedy: let a large model think and a small model tidy up.</p>
<p>The work arrives at a moment when the AI industry is pouring enormous resources into so-called agentic systems—models that autonomously plan tasks and invoke software tools on the user&#8217;s behalf. Tool invocation sits at the heart of this vision. In the standard tool-learning pipeline, which researchers typically divide into task planning, tool selection, tool invocation and response generation, the invocation stage is the make-or-break moment. The model must extract parameters from a natural-language query, match them to a tool&#8217;s specification, and emit a request so precisely structured that a downstream server can parse it without error. Any stray character, a missing parenthesis, or a misplaced comma can cause the entire call to fail silently.</p>
<p>What the researchers discovered through systematic perturbation experiments is striking: the success of a tool call is far more sensitive to format standardization than to semantic accuracy. When they fine-tuned the Llama3.1-8B-Instruct model on the ToolACE dataset using LoRA, randomly altering numbers in the training labels left accuracy nearly untouched, and shuffling parameter strings produced only a modest decline. But when they changed the format itself—swapping bracket types, converting integers to floating-point numbers, or reordering parameters—performance collapsed. Simply changing bracket styles dragged live-task accuracy down to 42.51 percent from a much higher baseline. Converting numbers to floats proved most devastating of all, with one metric plunging to 25.39 percent, because the abstract syntax tree evaluation used by the benchmark flags data-type mismatches instantly.</p>
<p>The most dramatic result came from compounding these perturbations. In a double mixed-modification experiment that first randomized bracket usage and then converted numbers to floats, live accuracy cratered to just 3.02 percent—essentially total failure. The lesson, the authors argue, is that conventional fine-tuning creates what they call format fragility: models rigidly cling to whatever format patterns they saw in training data, and even when prompts explicitly specify an output format, fine-tuned models frequently ignore those instructions and emit unparsable output. Earlier studies have described this phenomenon as format specialization or task locking, where intense fine-tuning erodes a model&#8217;s general in-context learning ability on non-target tasks.</p>
<p>Recognizing that reasoning and formatting are fundamentally different skills, the team designed a division-of-labor architecture that separates them. In their collaborative framework, the large language model receives the user&#8217;s question and a list of available tools, then produces an intermediate output containing its thought process, the selected tool name and the parameter information. Crucially, this intermediate output need not follow any format at all—the large model is freed from worrying about syntax. That freedom is precisely what preserves its generalization. The intermediate result is then handed to a small, specialized format model whose sole job is to normalize it into a strict, predefined structure that can be parsed directly into a callable API request.</p>
<p>The experimental payoff was substantial. When the same perturbed models were paired with the formatting model, accuracy rebounded dramatically. The Random Mix Twice configuration, which had fallen to 3.02 percent, soared to 73.42 percent once the small model normalized the output. The formatting step effectively absorbs all the chaotic variations—missing parentheses, wrong number formats, unexpected parameter orders—that would otherwise doom the invocation. The authors also contrast their approach with in-context learning, noting that few-shot examples struggle to exhaustively cover complex, nested parameter schemas, whereas a dedicated format model explicitly models the output structure and separates tool selection from argument generation.</p>
<p>The study used the Berkeley Function Call Leaderboard, a benchmark of more than 1,700 instances spanning simple, multiple, parallel and parallel-multiple function calls in Python, as well as REST API, JavaScript and Java tasks. Evaluation relied on the benchmark&#8217;s abstract syntax tree methodology, which checks whether function names, required parameters and data types all conform to the function documentation. The experiments ran on a single RTX 4090 GPU, underscoring that the collaborative method is computationally modest: instead of retraining a giant model, it attaches a lightweight normalizer to the end of the pipeline.</p>
<p>The implications reach across the AI industry. Giants including IBM&#8217;s Granite-20B-FunctionCalling, ToolLLM, APIGen and ToolACE have all pursued better function-calling models through increasingly sophisticated fine-tuning and dataset synthesis. But the new study suggests a quiet vulnerability running through that entire paradigm: as long as a single model is asked to be both reasoner and formatter, it will remain brittle. A wrong parameter value may still pass parsing and merely yield an irrelevant result, but a wrong bracket is fatal. The finding that format correctness outranks content accuracy inverts a common assumption that semantic quality is the primary axis of model quality, and it offers a practical, modular fix that developers could retrofit onto existing systems without touching the underlying model weights.</p>
<p>The authors are candid about the limits of their approach. Their evaluation remains confined to static, single-turn settings on specific test sets, while real-world agents must handle multi-turn dialogues that demand consistent tracking of context and parameters across turns. Real APIs also evolve their schemas over time, and format learning grounded in fixed training data struggles to adapt. Yet the multi-model framework points toward a natural solution: the large model can continue to handle contextual reasoning and intent understanding while the small model operates as a lightweight, updatable formatter that maps intent to whatever the current API schema requires. The team plans to integrate schema-based validation checkers and explore online adaptation techniques, including few-shot in-context learning and parameter-efficient fine-tuning, to keep inference costs low while boosting robustness in dynamic deployments.</p>
<p>For a field fixated on scale, the takeaway is refreshingly counterintuitive. Sometimes the fastest way to make a giant AI smarter is to pair it with a tiny, single-minded helper obsessed with punctuation. By decoupling what a model knows from how it says it, the study reframes tool learning as a coordination problem rather than a capability problem—and suggests that the next leap in autonomous AI agents may come not from bigger brains, but from better teamwork.</p>
<p><strong>Subject of Research:</strong> Enhanced tool invocation method through multi-model collaboration</p>
<p><strong>Article Title:</strong> Enhanced tool invocation method through multi-model collaboration</p>
<p><strong>Article References:</strong> Enhanced tool invocation method through multi-model collaboration. (n.d.). <a href="https://doi.org/10.1007/s44336-025-00028-7" rel="noopener noreferrer">https://doi.org/10.1007/s44336-025-00028-7</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1007/s44336-025-00028-7" rel="noopener noreferrer">10.1007/s44336-025-00028-7</a></p>
<p><strong>Keywords:</strong> Enhanced, tool, invocation, method, multi-model, collaboration, scientific research</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">205719</post-id>	</item>
		<item>
		<title>Free Water Imaging in Parkinson&#8217;s Disease Demands Methodological Nuance, Study Argues</title>
		<link>https://scienmag.com/free-water-imaging-in-parkinsons-disease-demands-methodological-nuance-study-argues/</link>
		
		<dc:creator><![CDATA[Diana Fleming]]></dc:creator>
		<pubDate>Sat, 12 Sep 2026 21:45:23 +0000</pubDate>
				<category><![CDATA[Medicine]]></category>
		<category><![CDATA[Biomarkers]]></category>
		<category><![CDATA[diffusion MRI]]></category>
		<category><![CDATA[diffusion-weighted MRI]]></category>
		<category><![CDATA[free water imaging]]></category>
		<category><![CDATA[free water imaging techniques]]></category>
		<category><![CDATA[image processing]]></category>
		<category><![CDATA[magnetic resonance imaging]]></category>
		<category><![CDATA[matters]]></category>
		<category><![CDATA[method]]></category>
		<category><![CDATA[methodological nuances in neuroimaging]]></category>
		<category><![CDATA[methodology]]></category>
		<category><![CDATA[MRI analytical methodology]]></category>
		<category><![CDATA[neurodegeneration]]></category>
		<category><![CDATA[neurodegeneration biomarkers]]></category>
		<category><![CDATA[neurodegeneration tracking]]></category>
		<category><![CDATA[neuroinflammation]]></category>
		<category><![CDATA[neuroinflammation detection]]></category>
		<category><![CDATA[Parkinson's disease]]></category>
		<category><![CDATA[Parkinson's disease diagnosis]]></category>
		<category><![CDATA[Parkinson's disease neuroimaging]]></category>
		<category><![CDATA[quantitative imaging markers]]></category>
		<category><![CDATA[substantia nigra]]></category>
		<category><![CDATA[substantia nigra neuronal loss]]></category>
		<category><![CDATA[tissue microstructure changes]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=198788</guid>

					<description><![CDATA[Researchers argue that free water imaging in Parkinson's disease produces method-dependent results that resist simple binary interpretation.]]></description>
										<content:encoded><![CDATA[<p>Free water imaging has become one of the most closely watched techniques in the effort to detect and track Parkinson&#8217;s disease with magnetic resonance imaging. The idea is elegantly simple: as neurons in the substantia nigra degenerate, the microscopic architecture of the tissue changes, and water molecules that once were constrained by cell membranes gain extra freedom to diffuse. By modeling this excess freely diffusing water, researchers hope to obtain a quantitative marker of neurodegeneration and, potentially, of the inflammatory processes that accompany it. A new commentary published in npj Parkinson&#8217;s Disease argues, however, that the field has too often treated the output of free water imaging as a straightforward verdict on disease, when in reality the measurement is deeply shaped by the analytical choices made along the way.</p>
<p>The technique rests on diffusion-weighted MRI, which sensitizes the MR signal to the random Brownian motion of water molecules. In a typical acquisition, the signal is measured along many diffusion-encoding directions, and a model is fitted to describe how the apparent diffusion coefficient varies with direction. In most brain tissue, diffusion is restricted and anisotropic, meaning water moves more easily along axonal bundles than across them. Free water imaging extends the standard diffusion tensor model by adding an isotropic compartment: a fraction of the voxel&#8217;s water is assumed to diffuse freely and equally in all directions, unconstrained by tissue microstructure. The estimated volume fraction of this compartment, often called the free water fraction, is the quantity that studies have linked to Parkinson&#8217;s disease.</p>
<p>What the commentary emphasizes is that this seemingly single number is, in practice, the product of a long chain of decisions. Every stage of the pipeline matters: the strength and number of diffusion-encoding gradients, the number of directions acquired, the echo time and voxel size, the correction for head motion and eddy currents, the approach to removing non-brain tissue, the handling of signal dropout, the fitting algorithm used to estimate the free water fraction, and the way regions of interest are defined in the midbrain. Each of these choices can shift the estimated values, and because different studies make different choices, their results are not always directly comparable.</p>
<p>This matters acutely in Parkinson&#8217;s disease research because the effect sizes involved are modest. The changes in free water fraction reported between people with Parkinson&#8217;s disease and healthy controls are typically small in absolute terms, often on the order of a few tenths of a percent to a few percent of the signal fraction. When the biological signal is that subtle, even small methodological differences can rival or exceed the effect being sought. A pipeline that smooths data aggressively, or that defines the substantia nigra generously, may report group differences where a more conservative pipeline finds none. Conversely, an underpowered or noisy acquisition may obscure real biology. The commentary&#8217;s central claim is that free water imaging findings in Parkinson&#8217;s disease should therefore be read as conditional statements, valid for a particular acquisition, preprocessing stream, and region-of-interest strategy, rather than as universal truths about the diseased brain.</p>
<p>The stakes are high because free water imaging has been proposed as a candidate imaging biomarker for disease progression and for use in clinical trials. Several longitudinal studies have suggested that free water fraction in the substantia nigra increases over time in people with Parkinson&#8217;s disease, raising hopes that the measure could serve as a sensitive endpoint for disease-modifying therapies. If those hopes are to be realized, the field needs to know how much of the measured change reflects biology and how much reflects the measurement apparatus. A biomarker that drifts with scanner software updates, or that responds more strongly to a change in preprocessing than to a change in the disease, cannot support the weight of a multi-center trial.</p>
<p>The commentary also addresses a conceptual trap: the tendency to interpret an elevated free water fraction as a direct, one-to-one readout of neuroinflammation. The biological rationale is plausible, because inflammatory processes such as astrocytic activation and microglial responses can expand the extracellular space and increase the mobility of water. But elevated free water is not specific to inflammation. Edema, enlarged perivascular spaces, tissue atrophy with partial volume effects from cerebrospinal fluid, and even residual artifacts from motion or susceptibility gradients can all inflate the estimate. Treating free water fraction as a binary indicator of an active inflammatory process, present or absent, oversimplifies what is in fact a composite measurement influenced by multiple tissue properties and multiple sources of error.</p>
<p>Partial volume contamination deserves particular attention in the midbrain, where the structures of interest are small and intimately surrounded by cerebrospinal fluid spaces. The substantia nigra lies adjacent to the interpeduncular cistern, and even with careful region-of-interest placement, signal from free cerebrospinal fluid can leak into the measured voxels, especially at the resolutions commonly used in research scanning. Some pipelines attempt to correct for this, while others rely on conservative masking. The commentary suggests that differences in how this problem is handled may explain a substantial portion of the variability in the literature, with some studies reporting robust group differences and others reporting null results for ostensibly similar comparisons.</p>
<p>None of this, the authors are careful to note, amounts to a dismissal of free water imaging. On the contrary, the technique remains one of the most promising MRI-based approaches to the nigral pathology that defines Parkinson&#8217;s disease, precisely because it targets a biologically meaningful property of tissue rather than a gross structural change that appears only late in the disease course. The argument is for methodological transparency and rigor: studies should report their acquisition parameters and preprocessing steps in full, share their analysis code where possible, and validate their pipelines against phantom data or across independent datasets. Harmonization efforts across scanning sites, and sensitivity analyses that show how results change under alternative processing choices, would allow the field to distinguish findings that are robust from those that are artifacts of a particular workflow.</p>
<p>For clinicians and trial designers, the practical message is one of calibrated expectations. Free water imaging is not yet a diagnostic test, and a single elevated value in an individual patient should not be read as a verdict on their disease state. The technique&#8217;s near-term value lies in group-level comparisons and longitudinal tracking within carefully controlled studies, where its sensitivity to change can be exploited while its methodological dependencies are held constant. As the field moves toward standardization, the commentary argues, the goal should be pipelines whose outputs are stable across sites and scanners, so that the biological signal of neurodegeneration can finally be separated from the technical noise of measurement. In free water imaging, the method is not a mere technicality; it is part of the result itself, and recognizing that is the first step toward turning an intriguing research measurement into a dependable clinical tool.</p>
<p><strong>Subject of Research:</strong> The influence of image processing methodology on free water imaging measurements in Parkinson&#x27;s disease</p>
<p><strong>Article Title:</strong> The method matters: free water imaging in Parkinson’s disease is not a binary verdict</p>
<p><strong>Article References:</strong> The method matters: free water imaging in Parkinson’s disease is not a binary verdict. (n.d.). <a href="https://doi.org/10.1038/s41531-026-01492-8" rel="noopener noreferrer">https://doi.org/10.1038/s41531-026-01492-8</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1038/s41531-026-01492-8" rel="noopener noreferrer">10.1038/s41531-026-01492-8</a></p>
<p><strong>Keywords:</strong> Parkinson&#x27;s disease, free water imaging, diffusion MRI, neuroinflammation, biomarkers, image processing, substantia nigra, magnetic resonance imaging, neurodegeneration, methodology, method, matters</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">198788</post-id>	</item>
	</channel>
</rss>
