<?xml version="1.0" encoding="UTF-8"?><rss version="2.0"
	xmlns:content="http://purl.org/rss/1.0/modules/content/"
	xmlns:wfw="http://wellformedweb.org/CommentAPI/"
	xmlns:dc="http://purl.org/dc/elements/1.1/"
	xmlns:atom="http://www.w3.org/2005/Atom"
	xmlns:sy="http://purl.org/rss/1.0/modules/syndication/"
	xmlns:slash="http://purl.org/rss/1.0/modules/slash/"
	>

<channel>
	<title>capability discovery &#8211; Science</title>
	<atom:link href="https://scienmag.com/tag/capability-discovery/feed/" rel="self" type="application/rss+xml" />
	<link>https://scienmag.com</link>
	<description></description>
	<lastBuildDate>Sun, 04 Oct 2026 04:26:23 +0000</lastBuildDate>
	<language>en-US</language>
	<sy:updatePeriod>
	hourly	</sy:updatePeriod>
	<sy:updateFrequency>
	1	</sy:updateFrequency>
	<generator>https://wordpress.org/?v=7.1.2</generator>

<image>
	<url>https://scienmag.com/wp-content/uploads/2024/07/cropped-scienmag_ico-32x32.jpg</url>
	<title>capability discovery &#8211; Science</title>
	<link>https://scienmag.com</link>
	<width>32</width>
	<height>32</height>
</image> 
<site xmlns="com-wordpress:feed-additions:1">73899611</site>	<item>
		<title>New Local Runtime Lets AI Agents Safely Command Your Desktop Files, Browser and Email</title>
		<link>https://scienmag.com/new-local-runtime-lets-ai-agents-safely-command-your-desktop-files-browser-and-email/</link>
		
		<dc:creator><![CDATA[Denise Maddox]]></dc:creator>
		<pubDate>Sun, 04 Oct 2026 04:26:23 +0000</pubDate>
				<category><![CDATA[Technology and Engineering]]></category>
		<category><![CDATA[AI agent desktop automation]]></category>
		<category><![CDATA[AI agents]]></category>
		<category><![CDATA[AI privacy protection measures]]></category>
		<category><![CDATA[AI-controlled file system access]]></category>
		<category><![CDATA[benchmark evaluation]]></category>
		<category><![CDATA[C# .NET 8 AI runtime]]></category>
		<category><![CDATA[capability discovery]]></category>
		<category><![CDATA[language model task planning and execution]]></category>
		<category><![CDATA[language models]]></category>
		<category><![CDATA[local runtime]]></category>
		<category><![CDATA[local runtime for AI command execution]]></category>
		<category><![CDATA[MIT licensed AI development tools]]></category>
		<category><![CDATA[open-source]]></category>
		<category><![CDATA[open-source AI workspace security]]></category>
		<category><![CDATA[OpenTangYuan]]></category>
		<category><![CDATA[openTangYuan software project]]></category>
		<category><![CDATA[research support]]></category>
		<category><![CDATA[responsible AI agent design]]></category>
		<category><![CDATA[safe AI browser and email integration]]></category>
		<category><![CDATA[software security]]></category>
		<category><![CDATA[trust boundary]]></category>
		<category><![CDATA[Windows automation]]></category>
		<category><![CDATA[Windows workstation AI capabilities]]></category>
		<category><![CDATA[workflow automation]]></category>
		<guid isPermaLink="false">https://scienmag.com/?p=233486</guid>

					<description><![CDATA[An open-source Windows runtime called OpenTangYuan lets AI language-model agents discover and execute local file, browser, email and tool operations behind a policy-controlled trust boundary, cutting workflow time and token use by roughly two-thirds in benchmark tests.]]></description>
										<content:encoded><![CDATA[<p>Artificial intelligence agents have become remarkably good at planning. Ask a modern language model to organize a pile of documents, gather evidence from a website, or draft a summary of recent correspondence, and it will happily produce a step-by-step strategy. What it cannot safely do, on its own, is reach into your computer and actually carry out those steps. Handing an external AI planner direct access to your mailbox, your file system, or your installed programs would be an open invitation to privacy disasters and irreversible side effects. A new open-source software project published in the journal SoftwareX tackles exactly this gap, offering a carefully fenced-off local runtime through which AI agents can discover and invoke the capabilities of a Windows workstation without ever being given the keys to the machine itself.</p>
<p>The system, called OpenTangYuan, was developed by Mingyang Liu, Wenjuan Wang, Jin Liu, Lele Ouyang, Hanyu Yang and Haifeng Pang and is released under the permissive MIT license. Its central design idea is a strict separation of responsibilities: the external language-model agent does the thinking, while a locally installed runtime does the touching. The runtime, written in C# on .NET 8, exposes REST entry points through which an agent can ask what the machine is capable of, retrieve the parameters of individual skills, and submit structured requests for execution. Every actual operation on files, browsers, email accounts, screenshots, or approved local programs happens inside the trusted local boundary, mediated by the runtime rather than by the agent.</p>
<p>Capability discovery is one of the system&#8217;s most elegant mechanisms. Each skill is described in a manifest file that lists its name, its typed parameters, and an invocation example. When an external agent wants to accomplish a task, it first calls a discovery endpoint that returns compact summaries of available workflows and built-in skills. If a stored workflow already matches the request, the agent receives its predefined steps; otherwise it retrieves the detailed schema of the individual skills it needs and composes a temporary multi-step workflow on the fly. This workflow-first approach means that routine, repetitive tasks can be captured once as reusable recipes, while genuinely novel requests can still be assembled from the same typed building blocks.</p>
<p>The built-in skill set reads like a menu of the most tedious chores in academic and office life: searching and organizing files, opening documents, capturing screenshots, driving a browser through Playwright automation, sending and reading email via MailKit, dispatching enterprise messaging notifications, invoking approved local analysis tools, and packaging results for delivery. Intermediate results flow through a shared execution context, so the file path discovered in step one can be reused by the copy operation in step three, and the screenshot captured in step five can be attached automatically by the email step in step six. Developers can extend the system by registering their own manifests and endpoints, turning the runtime into a growing library of locally controlled capabilities.</p>
<p>Security is where the design gets deliberately conservative. The authors are refreshingly candid about what their boundary does and does not protect. The implemented controls include API-key authentication on the agent-facing controller actions, required-argument validation, checks that file operations stay within allowed root directories, allowlists of approved executable names, and observable execution records. What the system explicitly does not provide is process sandboxing, semantic redaction of returned content, or full data-loss prevention. A permitted skill can still return derived text, file paths, or browser page content that crosses the trust boundary as structured output. The authors state plainly that the current release should not be interpreted as offering universal field-level filtering, and they warn that a success flag from the runtime does not by itself prove that an output is correct or free of sensitive content.</p>
<p>To demonstrate that the machinery actually works, the team built a reproducible benchmark of thirteen fixed cases using deterministic local fixtures. Seven cases tested functional and output correctness, judged against observable artifacts such as exact file paths, page output, and SHA-256 hash values. Three cases tested policy enforcement, confirming that invalid arguments, forbidden paths, and unapproved executables are correctly rejected. The remaining three probed workflow context passing and failure semantics, including a missing-source search followed by a failing copy and an unresolved template variable whose absence correctly propagates downstream. All thirteen cases conformed to their predefined oracles, and the distinction between transport status, runtime success, task correctness, and policy conformance proved essential: a rejected request can be the right outcome even though the task did not succeed.</p>
<p>A second evaluation, labeled RS01, exercised a genuine research-support workflow rather than an administrative stand-in. The runtime opened a specified website, retrieved the page title, captured a full-page screenshot into an allowed directory, invoked an allowlisted benchmark tool to compute the screenshot&#8217;s SHA-256 hash, and emailed the evidence with the screenshot attached. The hash was independently verified outside the workflow. Across three recorded runs plus one clean replay, all four executions satisfied every predefined check, from page access to attachment presence. The case shows how the same mechanisms that serve office automation can support research evidence collection with locally verifiable integrity.</p>
<p>The most striking quantitative result concerns the value of storing workflows rather than recomposing them from scratch each time. The team ran two fixed tasks, one for office document handling and one for research evidence gathering, in thirty paired comparisons each, with execution order varied to reduce ordering effects. Every one of the 120 executions passed its correctness checks, but the stored workflows were dramatically cheaper: median end-to-end time fell from roughly 49 seconds for ad-hoc composition to 15.5 seconds for stored workflows, and median token consumption dropped from about 24,500 to about 7,600. Wilcoxon signed-rank tests confirmed the reductions were statistically significant, with p-values below 0.001 for both time and tokens in both tasks. Median paired reductions reached 68 percent in time and 70 percent in tokens for one task, and 67.2 percent and 69.5 percent for the other.</p>
<p>Before the formal benchmark, the system also ran a historical pilot at the academic affairs office of the College of Arts and Information Engineering at Dalian Polytechnic University, deployed on a single Windows workstation in April 2026. Over its first ten days the runtime logged 1,061 executions across four task categories, with retained aggregate success rates of 97 percent for file search and copy, 99 percent for browser automation, 100 percent for email processing, and 98 percent for multi-step workflows. The authors are careful to note that these figures rest on runtime-reported status rather than an independent semantic oracle, and that the raw records are no longer retained, so the pilot stands as descriptive evidence of operational feasibility rather than proof of correctness.</p>
<p>The limitations are stated with unusual precision for a software release. The runtime is Windows-oriented, and full desktop functionality has not been evaluated on Linux or in containers. There is no multi-tenant isolation, no centralized cloud service, and no claim of protection against a compromised administrator, stolen credentials, or malicious behavior inside an already approved executable. Prompt injection and semantic exfiltration through permitted outputs remain outside the defended boundary. Yet within those limits, OpenTangYuan makes a compelling argument for a particular architectural pattern in the age of agentic AI: let the language model plan, but let an inspectable, policy-controlled local runtime execute. As institutions grapple with how to give AI assistants real utility without surrendering their data, this open-source execution boundary offers a concrete, benchmarked starting point that others can examine, extend, and deploy on their own terms.</p>
<p><strong>Subject of Research:</strong> A policy-controlled local runtime enabling AI agents to safely execute local research-support and office automation tasks on Windows</p>
<p><strong>Article Title:</strong> OpenTangYuan: A policy-controlled local runtime for agent-driven research support and office workflows</p>
<p><strong>Article References:</strong> Liu, M., Wang, W., Liu, J., Ouyang, L., Yang, H., &amp; Pang, H. (2026). OpenTangYuan: A policy-controlled local runtime for agent-driven research support and office workflows. <em>SoftwareX, 36</em>, Article 103097. <a href="https://doi.org/10.1016/j.softx.2026.103097" rel="noopener noreferrer">https://doi.org/10.1016/j.softx.2026.103097</a></p>
<p><strong>Image Credits:</strong> AI Generated</p>
<p><strong>DOI:</strong> <a href="https://doi.org/10.1016/j.softx.2026.103097" rel="noopener noreferrer">10.1016/j.softx.2026.103097</a></p>
<p><strong>Keywords:</strong> AI agents, OpenTangYuan, local runtime, workflow automation, capability discovery, software security, Windows automation, language models, research support, open source, benchmark evaluation, trust boundary</p>
]]></content:encoded>
					
		
		
		<post-id xmlns="com-wordpress:feed-additions:1">233486</post-id>	</item>
	</channel>
</rss>
