Wednesday, September 23, 2026
Science
No Result
View All Result
  • Login
  • HOME
  • SCIENCE NEWS
  • CONTACT US
  • HOME
  • SCIENCE NEWS
  • CONTACT US
No Result
View All Result
Scienmag
No Result
View All Result
Home Science News Technology and Engineering

Decision Trees for Giant Knowledge Bases: New Distributed Approach Scales Logical AI

September 23, 2026
in Technology and Engineering
Blake Davidson
By Blake Davidson Scienmag Editorial Profile - Data Science
Reading Time: 5 mins read
0
Decision Trees for Giant Knowledge Bases: New Distributed Approach Scales Logical AI

Decision Trees for Giant Knowledge Bases: New Distributed Approach Scales Logical AI

Decision Trees for Giant Knowledge Bases: New Distributed Approach Scales Logical AI

65
SHARES
587
VIEWS
Share on FacebookShare on Twitter
ADVERTISEMENT

Machine learning has grown accustomed to trading transparency for scale. The deep neural networks that dominate modern artificial intelligence can absorb billions of examples, but when they make a decision, even their creators often cannot say precisely why. A team of researchers led by Heba Mohamed of the University of Bonn and the University of Alexandria, working with Günter Kniesel-Wünsche, Jens Lehmann and Said Fathalla, has now tackled the opposite side of that trade-off: they have kept the logic fully transparent while pushing it to a scale that symbolic systems have never comfortably reached. Their new framework, called DistTDT, brings terminological decision tree learning, a classic of explainable machine learning, onto distributed computing clusters for the first time, and the results published in Knowledge and Information Systems suggest that white-box reasoning no longer has to stop at the memory limit of a single machine.

Terminological decision trees sit at the intersection of two venerable research traditions. On one side is the decision tree, the workhorse of classical machine learning, in which a cascade of yes-or-no tests sorts data into categories and can be read off directly by a human. On the other side is Description Logic, the family of formal languages that underlies the Semantic Web and the Web Ontology Language, known as OWL. In a standard decision tree, each node tests a simple feature. In a terminological decision tree, each node instead tests a logical concept, for example whether an individual belongs to the class of students who take at least one course. The tree thereby classifies individuals in a knowledge base and, in doing so, can propose new concept definitions that were never written down in the original ontology. That makes TDTs valuable not only for classification but for knowledge discovery, filling gaps in ontologies that encode domains from medicine to finance.

The catch has always been scale. Existing systems for inducing terminological decision trees, most notably the TermiTIS system and the ontology-driven decision tree approaches that followed it, run on a single machine and in memory. They were validated on ontologies with at most around 15,000 individuals. Meanwhile, the pressure to scale semantic machine learning has pushed much of the field toward sub-symbolic methods: graph neural networks, large language models and geometric ontology embeddings. These techniques distribute beautifully across clusters, but they sacrifice the strict logical guarantees that many applications cannot live without. Medical decision support, legal reasoning and regulatory compliance all demand classifiers whose reasoning can be audited step by step. Mohamed and her colleagues argue that DistTDT closes this gap, offering horizontal scalability without giving up symbolic exactness.

The architecture rests on Apache Spark, the distributed computing framework whose core abstraction, the Resilient Distributed Dataset, spreads partitioned data across worker nodes with fault tolerance. The workflow unfolds in three stages. First, the OWL layer of the SANSA framework converts the input ontology into distributed data structures, and a scalable statistics module extracts the classes, properties and individuals, partitioning them across the cluster. Second, a query generation phase produces artificial learning problems by combining two to eight primitive concepts with logical operators such as conjunction, union, negation and existential or universal restrictions. Third comes the actual tree induction, in which the driver coordinates worker nodes that evaluate candidate concepts in parallel against their local slices of instance data.

At the heart of the method lies a refinement operator, a device borrowed from inductive logic programming. Rather than randomly generating candidate concepts, as earlier approaches did, the operator systematically specializes or generalizes a parent concept to produce candidates that either subsume or are subsumed by it. Given the concept of a person who takes some course, the operator might produce a student who takes a semantic web course, or a graduate student who takes a course in both software engineering and artificial intelligence. Crucially, the operator exploits the subsumption hierarchy already encoded in the ontology’s terminology, which ensures that only refinements consistent with the existing knowledge base are generated. The number of possible specializations is in principle unlimited, so the system bounds the search breadth with a configurable parameter, preventing the combinatorial explosion from drowning the cluster in communication overhead.

Selecting which candidate concept to install as a new tree node relies on an information gain criterion computed with the Gini index, the same purity measure familiar from classical decision tree algorithms such as C4.5. Because some individuals may have uncertain membership in a concept, the authors extend the gain calculation to a ternary setting, treating positive, negative and uncertain individuals as separate impurity components. Once the best concept is chosen, the individuals are split into left and right subtrees, with uncertain individuals routed into both branches, and the algorithm recurses. A pruning threshold of 0.05 on the information gain keeps the trees from growing branches that add little statistical value, a choice the authors found balances accuracy against depth and cluster runtime.

Perhaps the most technically ambitious component is the prototype distributed reasoner that supports the whole enterprise. Inducing a terminological decision tree requires answering logical questions at every step: does this class have any instances, which individuals belong to it, and is a given assertion entailed by the knowledge base? Traditional tableau-based reasoners such as HermiT or the highly optimized Konclude are powerful but confined to a single machine, and their search spaces can grow exponentially. The new framework adopts a hybrid design. Heavy tableau mechanics, including backtracking and blocking to prevent infinite loops, run locally within each node’s memory. Above that, a distributed layer partitions logical dependencies, routes sub-queries across the cluster in parallel and aggregates the results. Candidate concepts travel to workers as broadcast variables, so the only network traffic is the stream of match counts flowing back to the driver. The authors validated the reasoner’s output with unit tests cross-checked by human domain experts.

The empirical evaluation, conducted on a five-node cluster of AMD Opteron machines with 64 cores and roughly 250 gigabytes of memory each, tested the framework on four expressive ontologies: Michalski’s classic trains problem, the medical MDM0.73 ontology, a financial ontology for e-banking and the mutagenesis benchmark. Against TermiTIS, the state-of-the-art centralized baseline, the distributed system dominated on every metric examined, including match rate, commission error rate, omission error rate and induction rate. The runtime comparison was stark. On the financial ontology, TermiTIS needed more than 340 minutes to complete its task, while DistTDT finished in under 40, roughly 11.8 percent of the baseline’s time. Consistency improved just as dramatically: on the medical dataset, the distributed system showed a standard deviation of 0.6 seconds against 72.8 for the centralized approach, with correspondingly narrower confidence intervals. The authors attribute much of this gain to the refinement operator, which replaces arbitrary concept generation with a structured traversal of the concept space.

Notably, the system sometimes classified individuals that a deductive reasoner could not decide at all, a phenomenon the authors measure as an induction rate. These induced classifications are logically new claims, not derivable from the ontology, but they may be exactly the kind of knowledge that helps populate incomplete knowledge bases, provided an ontology engineer reviews them. The team is candid about the trade-offs. Data skew from dense semantic hubs can stall tasks, though it does not affect correctness; Spark’s lineage-based fault tolerance can impose re-computation costs on node failures, mitigated here by selective caching; and distributing description logic syntax across Java virtual machines carries serialization overhead managed with the Kryo library. The approach also imposes some overhead on small, in-memory datasets where a single machine would suffice.

The researchers, whose implementation is open source and integrated into the SANSA framework as an extension built on Apache Spark, outline three directions for future work: extending the machinery beyond the current logic to more expressive OWL 2 profiles such as SROIQ, stress-testing against industrial-scale benchmarks exceeding millions of axioms, and building a dedicated benchmark for cluster scalability. For a field increasingly dominated by opaque statistical learners, the message of this work is a provocative one. Exact, auditable logical reasoning and planetary-scale data processing no longer have to be opposing choices, and the explainable branch of artificial intelligence may be ready for its own scaling moment.

Subject of Research: Distributed learning of terminological decision trees over large-scale OWL ontologies using Apache Spark and a distributed description logic reasoner.

Article Title: DistTDT: distributed terminological decision tree learning

Article References: Mohamed, H., Kniesel-Wünsche, G., Lehmann, J., & Fathalla, S. (2026). DistTDT: distributed terminological decision tree learning. Knowledge and Information Systems, 68(1), Article 263. https://doi.org/10.1007/s10115-026-02891-2

Image Credits: AI Generated

DOI: 10.1007/s10115-026-02891-2

Keywords: DistTDT, terminological decision trees, description logic, OWL, Apache Spark, distributed computing, concept learning, refinement operator, distributed reasoning, semantic web, explainable AI, SANSA framework

Cite Scienmag News

Blake Davidson. (September 23, 2026). Decision Trees for Giant Knowledge Bases: New Distributed Approach Scales Logical AI. Scienmag. https://scienmag.com/decision-trees-for-giant-knowledge-bases-new-distributed-approach-scales-logical-ai/

Blake Davidson. "Decision Trees for Giant Knowledge Bases: New Distributed Approach Scales Logical AI." Scienmag, 23 September 2026, https://scienmag.com/decision-trees-for-giant-knowledge-bases-new-distributed-approach-scales-logical-ai/. Accessed 23 September 2026.

Blake Davidson. "Decision Trees for Giant Knowledge Bases: New Distributed Approach Scales Logical AI." Scienmag. September 23, 2026. https://scienmag.com/decision-trees-for-giant-knowledge-bases-new-distributed-approach-scales-logical-ai/

Tags: Apache Sparkconcept learningdescription logicDescription Logic applicationsDistributed Computingdistributed computing for AIdistributed decision treesdistributed reasoningDistTDTexplainable AIExplainable Artificial Intelligenceexplainable machine learninglarge-scale knowledge baseslogic-based AI scalabilityOWLrefinement operatorSANSA frameworkscalable knowledge base reasoningsemantic websymbolic AI systemsterminological decision tree learningterminological decision treestransparent AI decision-makingwhite-box AI models
Share26Tweet16
Previous Post

Scientists Hunt Hidden Drought Genes in Ancient and Synthetic Wheats

Next Post

Hybrid AI With Attention Achieves Record Face Expression Recognition Accuracy

Related Posts

Hybrid AI With Attention Achieves Record Face Expression Recognition Accuracy
Technology and Engineering

Hybrid AI With Attention Achieves Record Face Expression Recognition Accuracy

September 23, 2026
Deep learning captures the heartbeat of the sarcomere in milliseconds
Technology and Engineering

Deep learning captures the heartbeat of the sarcomere in milliseconds

September 23, 2026
Designer Salts Called Ionic Liquids Are Reshaping How Metal Nanoparticles Are Made and Used in Medicine
Technology and Engineering

Designer Salts Called Ionic Liquids Are Reshaping How Metal Nanoparticles Are Made and Used in Medicine

September 23, 2026
TinZr: A $65 Open-Source Wearable Puts Multi-Sensor Health Tracking on a Single Tiny Board
Technology and Engineering

TinZr: A $65 Open-Source Wearable Puts Multi-Sensor Health Tracking on a Single Tiny Board

September 23, 2026
Tensor-Powered Receiver Promises More Reliable Drone Communications in Chaotic Airwaves
Technology and Engineering

Tensor-Powered Receiver Promises More Reliable Drone Communications in Chaotic Airwaves

September 23, 2026
Kagome Antiferromagnet Shatters Records for Magnetic-Field-Switched Hall Conductivity
Technology and Engineering

Kagome Antiferromagnet Shatters Records for Magnetic-Field-Switched Hall Conductivity

September 23, 2026
Next Post
Hybrid AI With Attention Achieves Record Face Expression Recognition Accuracy

Hybrid AI With Attention Achieves Record Face Expression Recognition Accuracy

  • Mothers who receive childcare support from maternal grandparents show more optimized

    Mothers who receive childcare support from maternal grandparents show more parental warmth, finds NTU Singapore study

    27656 shares
    Share 11059 Tweet 6912
  • University of Seville Breaks 120-Year-Old Mystery, Revises a Key Einstein Concept

    1061 shares
    Share 424 Tweet 265
  • Bee body mass, pathogens and local climate influence heat tolerance

    682 shares
    Share 273 Tweet 171
  • Researchers record first-ever images and data of a shark experiencing a boat strike

    546 shares
    Share 218 Tweet 137
  • Groundbreaking Clinical Trial Reveals Lubiprostone Enhances Kidney Function

    531 shares
    Share 212 Tweet 133
Science

Embark on a thrilling journey of discovery with Scienmag.com—your ultimate source for cutting-edge breakthroughs. Immerse yourself in a world where curiosity knows no limits and tomorrow’s possibilities become today’s reality!

RECENT NEWS

  • Hybrid AI With Attention Achieves Record Face Expression Recognition Accuracy
  • Decision Trees for Giant Knowledge Bases: New Distributed Approach Scales Logical AI
  • Scientists Hunt Hidden Drought Genes in Ancient and Synthetic Wheats
  • Magnetic fingerprints and AI reveal how traffic pollution hides in city soils

Categories

  • Agriculture
  • Anthropology
  • Archaeology
  • Athmospheric
  • Biology
  • Biotechnology
  • Blog
  • Bussines
  • Cancer
  • Chemistry
  • Climate
  • Earth Science
  • Editorial Policy
  • Marine
  • Mathematics
  • Medicine
  • Pediatry
  • Policy
  • Psychology & Psychiatry
  • Science Education
  • Social Science
  • Space
  • Technology and Engineering

Subscribe to Blog via Email

Enter your email address to subscribe to this blog and receive notifications of new posts by email.

Join 5,151 other subscribers

© 2025 Scienmag - Science Magazine

Welcome Back!

Login to your account below

Forgotten Password?

Retrieve your password

Please enter your username or email address to reset your password.

Log In
No Result
View All Result
  • HOME
  • SCIENCE NEWS
  • CONTACT US

© 2025 Scienmag - Science Magazine

Discover more from Science

Subscribe now to keep reading and get access to the full archive.

Continue reading