Sunday, October 4, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

New Local Runtime Lets AI Agents Safely Command Your Desktop Files, Browser and Email

October 4, 2026
in Technology and Engineering
Denise Maddox
By Denise Maddox Scienmag Editorial Profile - Mechanical Engineering
Reading Time: 5 mins read
0
New Local Runtime Lets AI Agents Safely Command Your Desktop Files, Browser and Email

New Local Runtime Lets AI Agents Safely Command Your Desktop Files, Browser and Email

New Local Runtime Lets AI Agents Safely Command Your Desktop Files, Browser and Email

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Artificial intelligence agents have become remarkably good at planning. Ask a modern language model to organize a pile of documents, gather evidence from a website, or draft a summary of recent correspondence, and it will happily produce a step-by-step strategy. What it cannot safely do, on its own, is reach into your computer and actually carry out those steps. Handing an external AI planner direct access to your mailbox, your file system, or your installed programs would be an open invitation to privacy disasters and irreversible side effects. A new open-source software project published in the journal SoftwareX tackles exactly this gap, offering a carefully fenced-off local runtime through which AI agents can discover and invoke the capabilities of a Windows workstation without ever being given the keys to the machine itself.

The system, called OpenTangYuan, was developed by Mingyang Liu, Wenjuan Wang, Jin Liu, Lele Ouyang, Hanyu Yang and Haifeng Pang and is released under the permissive MIT license. Its central design idea is a strict separation of responsibilities: the external language-model agent does the thinking, while a locally installed runtime does the touching. The runtime, written in C# on .NET 8, exposes REST entry points through which an agent can ask what the machine is capable of, retrieve the parameters of individual skills, and submit structured requests for execution. Every actual operation on files, browsers, email accounts, screenshots, or approved local programs happens inside the trusted local boundary, mediated by the runtime rather than by the agent.

Capability discovery is one of the system’s most elegant mechanisms. Each skill is described in a manifest file that lists its name, its typed parameters, and an invocation example. When an external agent wants to accomplish a task, it first calls a discovery endpoint that returns compact summaries of available workflows and built-in skills. If a stored workflow already matches the request, the agent receives its predefined steps; otherwise it retrieves the detailed schema of the individual skills it needs and composes a temporary multi-step workflow on the fly. This workflow-first approach means that routine, repetitive tasks can be captured once as reusable recipes, while genuinely novel requests can still be assembled from the same typed building blocks.

The built-in skill set reads like a menu of the most tedious chores in academic and office life: searching and organizing files, opening documents, capturing screenshots, driving a browser through Playwright automation, sending and reading email via MailKit, dispatching enterprise messaging notifications, invoking approved local analysis tools, and packaging results for delivery. Intermediate results flow through a shared execution context, so the file path discovered in step one can be reused by the copy operation in step three, and the screenshot captured in step five can be attached automatically by the email step in step six. Developers can extend the system by registering their own manifests and endpoints, turning the runtime into a growing library of locally controlled capabilities.

Security is where the design gets deliberately conservative. The authors are refreshingly candid about what their boundary does and does not protect. The implemented controls include API-key authentication on the agent-facing controller actions, required-argument validation, checks that file operations stay within allowed root directories, allowlists of approved executable names, and observable execution records. What the system explicitly does not provide is process sandboxing, semantic redaction of returned content, or full data-loss prevention. A permitted skill can still return derived text, file paths, or browser page content that crosses the trust boundary as structured output. The authors state plainly that the current release should not be interpreted as offering universal field-level filtering, and they warn that a success flag from the runtime does not by itself prove that an output is correct or free of sensitive content.

To demonstrate that the machinery actually works, the team built a reproducible benchmark of thirteen fixed cases using deterministic local fixtures. Seven cases tested functional and output correctness, judged against observable artifacts such as exact file paths, page output, and SHA-256 hash values. Three cases tested policy enforcement, confirming that invalid arguments, forbidden paths, and unapproved executables are correctly rejected. The remaining three probed workflow context passing and failure semantics, including a missing-source search followed by a failing copy and an unresolved template variable whose absence correctly propagates downstream. All thirteen cases conformed to their predefined oracles, and the distinction between transport status, runtime success, task correctness, and policy conformance proved essential: a rejected request can be the right outcome even though the task did not succeed.

A second evaluation, labeled RS01, exercised a genuine research-support workflow rather than an administrative stand-in. The runtime opened a specified website, retrieved the page title, captured a full-page screenshot into an allowed directory, invoked an allowlisted benchmark tool to compute the screenshot’s SHA-256 hash, and emailed the evidence with the screenshot attached. The hash was independently verified outside the workflow. Across three recorded runs plus one clean replay, all four executions satisfied every predefined check, from page access to attachment presence. The case shows how the same mechanisms that serve office automation can support research evidence collection with locally verifiable integrity.

The most striking quantitative result concerns the value of storing workflows rather than recomposing them from scratch each time. The team ran two fixed tasks, one for office document handling and one for research evidence gathering, in thirty paired comparisons each, with execution order varied to reduce ordering effects. Every one of the 120 executions passed its correctness checks, but the stored workflows were dramatically cheaper: median end-to-end time fell from roughly 49 seconds for ad-hoc composition to 15.5 seconds for stored workflows, and median token consumption dropped from about 24,500 to about 7,600. Wilcoxon signed-rank tests confirmed the reductions were statistically significant, with p-values below 0.001 for both time and tokens in both tasks. Median paired reductions reached 68 percent in time and 70 percent in tokens for one task, and 67.2 percent and 69.5 percent for the other.

Before the formal benchmark, the system also ran a historical pilot at the academic affairs office of the College of Arts and Information Engineering at Dalian Polytechnic University, deployed on a single Windows workstation in April 2026. Over its first ten days the runtime logged 1,061 executions across four task categories, with retained aggregate success rates of 97 percent for file search and copy, 99 percent for browser automation, 100 percent for email processing, and 98 percent for multi-step workflows. The authors are careful to note that these figures rest on runtime-reported status rather than an independent semantic oracle, and that the raw records are no longer retained, so the pilot stands as descriptive evidence of operational feasibility rather than proof of correctness.

The limitations are stated with unusual precision for a software release. The runtime is Windows-oriented, and full desktop functionality has not been evaluated on Linux or in containers. There is no multi-tenant isolation, no centralized cloud service, and no claim of protection against a compromised administrator, stolen credentials, or malicious behavior inside an already approved executable. Prompt injection and semantic exfiltration through permitted outputs remain outside the defended boundary. Yet within those limits, OpenTangYuan makes a compelling argument for a particular architectural pattern in the age of agentic AI: let the language model plan, but let an inspectable, policy-controlled local runtime execute. As institutions grapple with how to give AI assistants real utility without surrendering their data, this open-source execution boundary offers a concrete, benchmarked starting point that others can examine, extend, and deploy on their own terms.

Subject of Research: A policy-controlled local runtime enabling AI agents to safely execute local research-support and office automation tasks on Windows

Article Title: OpenTangYuan: A policy-controlled local runtime for agent-driven research support and office workflows

Article References: Liu, M., Wang, W., Liu, J., Ouyang, L., Yang, H., & Pang, H. (2026). OpenTangYuan: A policy-controlled local runtime for agent-driven research support and office workflows. SoftwareX, 36, Article 103097. https://doi.org/10.1016/j.softx.2026.103097

Image Credits: AI Generated

DOI: 10.1016/j.softx.2026.103097

Keywords: AI agents, OpenTangYuan, local runtime, workflow automation, capability discovery, software security, Windows automation, language models, research support, open source, benchmark evaluation, trust boundary

Cite Scienmag News

Denise Maddox. (October 4, 2026). New Local Runtime Lets AI Agents Safely Command Your Desktop Files, Browser and Email. Scienmag. https://scienmag.com/new-local-runtime-lets-ai-agents-safely-command-your-desktop-files-browser-and-email/

Denise Maddox. "New Local Runtime Lets AI Agents Safely Command Your Desktop Files, Browser and Email." Scienmag, 4 October 2026, https://scienmag.com/new-local-runtime-lets-ai-agents-safely-command-your-desktop-files-browser-and-email/. Accessed 4 October 2026.

Denise Maddox. "New Local Runtime Lets AI Agents Safely Command Your Desktop Files, Browser and Email." Scienmag. October 4, 2026. https://scienmag.com/new-local-runtime-lets-ai-agents-safely-command-your-desktop-files-browser-and-email/

Tags: AI agent desktop automationAI agentsAI privacy protection measuresAI-controlled file system accessbenchmark evaluationC# .NET 8 AI runtimecapability discoverylanguage model task planning and executionlanguage modelslocal runtimelocal runtime for AI command executionMIT licensed AI development toolsopen-sourceopen-source AI workspace securityOpenTangYuanopenTangYuan software projectresearch supportresponsible AI agent designsafe AI browser and email integrationsoftware securitytrust boundaryWindows automationWindows workstation AI capabilitiesworkflow automation
Share26Tweet16
Previous Post

One Truck a Day: Postal Policy Change Could Reject Tens of Thousands of Mail Ballots

Next Post

Machine Learning Meets Classical Statistics to Catch the Subtle Fingerprints of Polygenic Adaptation

Related Posts

Kitchen Blender Beats Ball Milling in Greener Route to Superstrong Conductive Nanocomposites
Technology and Engineering

Kitchen Blender Beats Ball Milling in Greener Route to Superstrong Conductive Nanocomposites

October 4, 2026
Breathing Liquid: Lung Volume Holds the Key to Safer Total Liquid Ventilation in Newborns
Technology and Engineering

Breathing Liquid: Lung Volume Holds the Key to Safer Total Liquid Ventilation in Newborns

October 4, 2026
Deep Mines Run Hot: Why 40°C Makes Coal Waste Concrete Stronger, Then Weaker
Technology and Engineering

Deep Mines Run Hot: Why 40°C Makes Coal Waste Concrete Stronger, Then Weaker

October 4, 2026
New JMIR Cardio Section Seeks Research on Generative and Multimodal AI in Heart Care
Technology and Engineering

New JMIR Cardio Section Seeks Research on Generative and Multimodal AI in Heart Care

October 4, 2026
Smart MRI Nanoprobes Switch On Inside Tumors to Reveal Cancer and Immune Battles
Technology and Engineering

Smart MRI Nanoprobes Switch On Inside Tumors to Reveal Cancer and Immune Battles

October 4, 2026
AI Spots Grape Leaf Diseases With Over 99 Percent Accuracy
Technology and Engineering

AI Spots Grape Leaf Diseases With Over 99 Percent Accuracy

October 4, 2026
Next Post
Machine Learning Meets Classical Statistics to Catch the Subtle Fingerprints of Polygenic Adaptation

Machine Learning Meets Classical Statistics to Catch the Subtle Fingerprints of Polygenic Adaptation

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Machine Learning Meets Classical Statistics to Catch the Subtle Fingerprints of Polygenic Adaptation
  • New Local Runtime Lets AI Agents Safely Command Your Desktop Files, Browser and Email
  • One Truck a Day: Postal Policy Change Could Reject Tens of Thousands of Mail Ballots
  • Tiny Soil Worms Rewrite the Rules of Life on Arid Mountains

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,149 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading