Thursday, August 6, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

NSF CAREER Award Supports Illinois Researcher’s Effort to Improve AI Training Data

August 6, 2026
in Technology and Engineering
Reading Time: 3 mins read
0
NSF CAREER Award Supports Illinois Researcher’s Effort to Improve AI Training Data

NSF CAREER Award Supports Illinois Researcher’s Effort to Improve AI Training Data

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Assistant Professor Jiaqi Ma has received a five-year, $660,307 National Science Foundation CAREER award to investigate one of the most consequential and least understood forces shaping modern artificial intelligence: the data used to train it. His project, “Data Attribution and Curation for Web-Scale AI Systems,” will develop methods for determining how individual training examples influence the behavior of large AI models, potentially changing how researchers build, audit, improve, and govern these systems.

The NSF CAREER award is among the most selective honors available to early-career researchers in the United States. It recognizes faculty members whose work combines strong research potential with a commitment to education and broader public impact. Ma, an assistant professor at the University of Illinois School of Information Sciences, will use the award to study data attribution, a family of techniques designed to estimate the contribution of specific training examples to a model’s predictions, capabilities, errors, or harmful behavior.

The challenge is enormous because today’s AI systems are trained on datasets that can contain billions or even trillions of pieces of information. These collections may include books, websites, images, code, conversations, scientific papers, and other forms of digital content gathered from many sources. Training data can be duplicated, outdated, biased, misleading, copyrighted, or difficult to trace. Once these examples have been processed through a model’s training procedure, identifying which data influenced a particular output becomes extremely difficult.

Data attribution seeks to make that influence measurable. In principle, an attribution method could help determine whether a model’s answer was supported by reliable information, shaped by a particular document, or affected by a problematic example. It could also help researchers identify which training samples improve performance on a task and which ones cause errors. Such information would provide a more precise alternative to treating an entire dataset as equally valuable.

“Data is one of the most important ingredients in modern AI,” Ma said, “but we still have a limited understanding of how individual training examples shape a model’s behavior.” His project will focus on developing attribution tools that remain useful at web scale, where the size and constant evolution of datasets make conventional analysis difficult. The research will also address how attribution can work when models are repeatedly updated with new information rather than trained only once.

Understanding these relationships could transform data curation, the process of deciding what information should be included in a training dataset. Instead of relying primarily on broad filtering rules or simple quality scores, developers could use attribution signals to prioritize examples that improve reasoning, remove data associated with recurring failures, and identify content that contributes to unsafe or unreliable outputs. The same approach could help reveal gaps in a dataset, including areas where a model performs poorly because relevant perspectives or information are missing.

The implications extend beyond accuracy. AI systems increasingly influence search results, recommendations, education, employment, scientific research, and creative work. If training data contains harmful stereotypes, misinformation, or material that should not have been used, those problems can be reflected in model behavior. Attribution could offer researchers a way to trace such effects back to their sources and evaluate whether removing or modifying particular examples changes the system’s responses. It may also support machine unlearning, in which a model is adjusted to reduce the influence of selected data after training.

Ma also sees a possible connection between data attribution and compensation. Much of the content used to develop AI systems is created by people whose contributions may be difficult to identify once the material has been absorbed into a massive dataset. More transparent attribution methods could eventually help determine whether and how particular works contribute to a model’s capabilities. That could create a technical foundation for recognizing contributors and designing more equitable systems for compensating them, although the legal and economic mechanisms for doing so remain unresolved.

The project will combine research with education and public access. Ma plans to create open-source software, conference tutorials, course modules, and other educational resources. The effort will provide research opportunities for graduate and undergraduate students, as well as pre-college learners, with the goal of widening participation in data-centered AI research. By making the tools publicly available, the project could allow academic researchers, technology developers, and policymakers to examine AI systems using shared methods rather than relying only on proprietary evaluations.

Ma’s research spans machine learning and artificial intelligence, with a particular focus on the data foundations of AI. His work examines how training data affect models through attribution, how data-centric algorithms can improve the quality and safety of datasets through curation and synthetic data generation, and how data mediate the societal consequences of AI through compensation and machine unlearning. He received a Best Paper Award at the 2024 International Conference on Learning Representations workshop on navigating and addressing data problems for foundation models and was named a New Faculty Highlight by the Association for the Advancement of Artificial Intelligence in 2025. Ma earned his PhD from the University of Michigan and completed postdoctoral research at Harvard University.

Subject of Research: Data attribution and curation for web-scale artificial intelligence systems

Keywords

Artificial intelligence, machine learning, data attribution, data curation, training data, large language models, foundation models, machine unlearning, synthetic data, AI safety, AI transparency, NSF CAREER award

Tags: AI training data analysischallenges in managing massive training datasetsdata attribution in machine learningdata curation for web-scale AIearly-career AI researcher Illinoisenhancing AI model accuracy through data attributionethical governance of AI training datasetsimpact of training data on AI model behaviorimproving AI model transparency and accountabilitylarge-scale AI system developmentmethods for training data influence estimationNSF CAREER award for AI research
Share26Tweet16
Previous Post

Collaborative team doubles patient-derived in vitro cancer models available for research

Next Post

Membrane-Disrupting Peptide Triggers Immune-Stimulating Cancer Cell Death

Related Posts

Engineers develop AI tool to manage tomorrow’s complex power grid
Technology and Engineering

Engineers develop AI tool to manage tomorrow’s complex power grid

August 6, 2026
Membrane-Disrupting Peptide Triggers Immune-Stimulating Cancer Cell Death
Medicine

Membrane-Disrupting Peptide Triggers Immune-Stimulating Cancer Cell Death

August 6, 2026
No-till and microbial fertilizers jointly boost carbon storage in nutrient-poor albic soils
Technology and Engineering

No-till and microbial fertilizers jointly boost carbon storage in nutrient-poor albic soils

August 6, 2026
Quantum Simulator Reveals Pseudogap in Fermi–Hubbard Model
Medicine

Quantum Simulator Reveals Pseudogap in Fermi–Hubbard Model

August 6, 2026
KAIST Reconstructs Transparent Objects Through Dynamic Scattering Layers in One Shot
Technology and Engineering

KAIST Reconstructs Transparent Objects Through Dynamic Scattering Layers in One Shot

August 6, 2026
Revealing How Female Restitution Occurs in Sugarcane Hybrids
Medicine

Revealing How Female Restitution Occurs in Sugarcane Hybrids

August 5, 2026
Next Post
Membrane-Disrupting Peptide Triggers Immune-Stimulating Cancer Cell Death

Membrane-Disrupting Peptide Triggers Immune-Stimulating Cancer Cell Death

  • Mothers who receive childcare support from maternal grandparents show more

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Scientists discover tiny vortices swirling across the Sun’s surface
  • Barrow Institute Receives $5 Million for Ibogaine Traumatic Brain Injury Trial
  • How Vaccine Conspiracy Beliefs Form and Spread
  • Human cancer models speed research and advance precision therapies

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Success! An email was just sent to confirm your subscription. Please find the email now and click 'Confirm Follow' to start subscribing.

Join 5,149 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine