A team of computer scientists at Princess Sumaya University for Technology in Amman, Jordan, has unveiled a hybrid artificial intelligence system designed to answer one of the most pressing questions in modern publishing: was this scientific text written by a human or by a machine? The new tool, called the Artificial Intelligence Scientific Text Detector, or AISciDetector, was developed by researchers Arwa Bader and Bushra Alhijawi and described in the journal Neural Computing and Applications. It combines a transformer-based text encoder with a deep neural network architecture to distinguish human-written scientific passages from text produced by large language models such as ChatGPT, Gemini, and LLaMA-3. In benchmark experiments, the system achieved an accuracy of 91.5 percent and an F1-score of 91.49 percent, outperforming both classical machine learning baselines and several transformer-based competitors tested by the authors.
The motivation behind the work is straightforward. Large language models have become remarkably fluent generators of scientific prose, capable of producing abstracts, introductions, and literature reviews that read convincingly like human scholarship. That fluency brings real benefits for drafting and editing, but it also opens the door to plagiarism, misinformation, and academic dishonesty at unprecedented scale. Journals, universities, and funding agencies now face a flood of submissions in which machine-generated passages may be indistinguishable to human readers. Previous studies have shown that even trained reviewers can be fooled: research comparing ChatGPT-generated abstracts with real ones found that blinded human reviewers often struggled to tell them apart, and commercial detection tools have been documented mislabeling genuine human manuscripts as AI creations. A reliable, scientifically focused detector has therefore become a matter of protecting the integrity and credibility of the academic record itself.
The technical heart of AISciDetector is a two-stage pipeline that marries two complementary ways of understanding text. In the first stage, the system uses BigBird, a transformer architecture designed to handle long documents efficiently. Standard transformer models such as BERT process text with a mechanism called self-attention, in which every word can, in principle, attend to every other word, but this becomes computationally expensive as documents grow longer. BigBird sidesteps that bottleneck with a sparse attention scheme, allowing the model to capture long-range dependencies, the distant relationships between words and concepts that are especially important in scientific writing, where a conclusion at the end of a passage may hinge on a definition or result stated paragraphs earlier. These BigBird embeddings serve as rich numerical representations of each text, encoding both its vocabulary and its broader structural patterns.
In the second stage, those embeddings are fed into a hybrid CNN-BiLSTM framework. The convolutional neural network, or CNN, acts as a local feature extractor: it slides filters across the embedding sequences to pick up short, characteristic patterns, much as convolutional networks detect edges and textures in images. Immediately after, a bidirectional long short-term memory network, or BiLSTM, models the sequential structure of the text in both directions. LSTMs are a form of recurrent neural network engineered to remember information over long sequences while avoiding the vanishing-gradient problems that plague simpler recurrent models. By running two LSTMs, one reading the sequence forward and one reading it backward, the BiLSTM layer builds a representation that reflects both what precedes and what follows each point in the text. The combined features then drive a binary classifier that outputs a single decision: human-written or LLM-generated.
Training and evaluating such a system requires data, and here the researchers made a second significant contribution. Alongside the detector, they introduce LLMSciTxt, a new dataset containing scientific texts written by human authors alongside outputs generated by three of the most widely used large language models: ChatGPT, Gemini, and LLaMA-3. This multi-model design matters because detectors trained on the output of a single generator often fail when confronted with text from a different one. By including outputs from three distinct model families, each with its own training data, architecture, and stylistic quirks, the dataset encourages the detector to learn general signatures of machine generation rather than the fingerprint of one particular system. The authors state that the datasets generated and analyzed during the study are available from the corresponding author on reasonable request.
The experimental results place AISciDetector ahead of the field it was tested against. The 91.5 percent accuracy and 91.49 percent F1-score, a metric that balances precision and recall and is particularly informative when the two error types carry different costs, exceeded the performance of classical machine learning approaches as well as transformer-based models evaluated under the same conditions. The authors attribute the gains to the hybrid design: BigBird supplies the long-context understanding that shallower models lack, while the CNN-BiLSTM stack adds a layer of local and sequential pattern recognition on top of those embeddings. In effect, the system looks at machine-generated text through three lenses at once, global context, local phrasing, and temporal flow, and the combination appears to capture subtle statistical traces that machines leave behind even when their prose reads naturally.
The study also situates itself within a rapidly growing research landscape. Earlier work by the same group explored deep learning detection methods for LLM-generated scientific content and a transformer-based approach presented at a computing sciences conference, and the new paper builds directly on that lineage. Around the world, other teams have pursued parallel strategies: some detectors analyze token probability sequences without any training, others frame human text as an out-of-distribution anomaly, and still others apply adversarial learning or statistical guarantees to make detection more robust. Benchmarks such as large-scale datasets of ChatGPT-written abstracts and real-world detection challenges have emerged to standardize evaluation. Yet the field remains contested, with independent evaluations of commercial tools such as GPTZero, ZeroGPT, Copyleaks, and Scribbr’s detector finding inconsistent reliability, and documented cases of human-written scientific manuscripts being wrongly flagged as machine-generated, a false positive that can carry serious consequences for authors.
That tension between sensitivity and fairness is precisely why domain-specific detectors like AISciDetector may prove valuable. Scientific text has distinctive properties, dense terminology, formulaic section structures, and heavy citation conventions, that generic detectors trained on news or student essays may handle poorly. By training and testing specifically on scientific content from multiple generators, the Jordanian team’s system targets the exact domain where the stakes are highest: peer-reviewed publishing. The authors argue that effective identification of LLM-generated content supports the integrity and credibility of academic publications, a goal shared by editors confronting a rising volume of submissions in which the boundary between human and machine authorship has blurred. The work also connects to broader efforts in computational linguistics, literature mining, and the emerging science of explainability for large language models, all of which seek to make machine-generated language more transparent and accountable.
Challenges remain before such tools can be deployed as gatekeepers. Language models evolve quickly, and each new generation of generators may produce text whose statistical fingerprints differ from those in today’s training data, meaning detectors must be continually retrained and revalidated. Cross-domain performance, robustness to paraphrasing or deliberate evasion, and the ethical weight of false accusations all demand careful handling, and the authors themselves note that detection and anti-detection exist in a continuing arms race. Still, the study offers a concrete demonstration that combining long-range transformer embeddings with bidirectional sequence modeling can push detection accuracy above 91 percent on scientific text spanning three major model families. As universities and journals grapple with how to police AI-assisted writing without punishing honest authors, tools like AISciDetector represent a step toward detection methods that are both technically sophisticated and tailored to the specific texture of scientific prose, helping ensure that the literature of science remains a record of genuine human inquiry, however much machines may assist along the way.
Subject of Research: Detection of large language model-generated scientific text using a hybrid BigBird and CNN-BiLSTM deep learning model
Article Title: LLM-generated scientific content detection method
Article References: Bader, A., & Alhijawi, B. (2026). LLM-generated scientific content detection method. Neural Computing and Applications, 38(17), Article 728. https://doi.org/10.1007/s00521-026-12465-6
Image Credits: AI Generated
DOI: 10.1007/s00521-026-12465-6
Keywords: large language models, AI text detection, scientific integrity, BigBird, CNN-BiLSTM, ChatGPT, Gemini, LLaMA-3, academic plagiarism, deep learning, natural language processing, transformers
Cite Scienmag News
Blake Davidson. (October 5, 2026). New AI Detector Spots Machine-Written Science Papers With 91.5% Accuracy. Scienmag. https://scienmag.com/new-ai-detector-spots-machine-written-science-papers-with-91-5-accuracy/
Blake Davidson. "New AI Detector Spots Machine-Written Science Papers With 91.5% Accuracy." Scienmag, 5 October 2026, https://scienmag.com/new-ai-detector-spots-machine-written-science-papers-with-91-5-accuracy/. Accessed 5 October 2026.
Blake Davidson. "New AI Detector Spots Machine-Written Science Papers With 91.5% Accuracy." Scienmag. October 5, 2026. https://scienmag.com/new-ai-detector-spots-machine-written-science-papers-with-91-5-accuracy/

