Community question answering platforms such as Stack Overflow have become indispensable infrastructure for the modern knowledge economy, but their success hinges on a deceptively simple operation: getting the right question in front of the right person. As questions grow more complex, a single expert is often no longer enough, and platforms increasingly need to assemble small teams whose combined knowledge can resolve a problem quickly and accurately. A new study published in the journal Machine Learning by Roohollah Etemadi, Morteza Zihayat, Kuan Feng, Jason Adelman, Fattane Zarrinkalam and Ebrahim Bagheri tackles this challenge head-on, presenting a neural framework that learns to discover both individual experts and collaborative teams of experts in a single, end-to-end trainable model.
The core limitation the researchers identify is conceptual rather than computational. Existing expert-finding methods are built to score the suitability of one expert for one question, treating each user as an isolated candidate. That design works reasonably well when a question falls squarely within a single specialty, but it breaks down when a question spans multiple domains or requires complementary skills. The team formation methods that do exist, meanwhile, tend to rely on heuristic search procedures that are computationally expensive and disconnected from the learning of the representations that describe experts in the first place. The new approach bridges this gap by jointly learning how to represent experts, questions and their interactions within one unified architecture.
Technically, the framework rests on two complementary encoders. The first is a structure encoder built on graph convolutional networks, the now-standard family of neural architectures for learning over graph-structured data. The community question answering platform is modeled as a graph in which users, questions and answers form interconnected nodes, and the graph convolutional network aggregates information from a node’s neighbors to produce an embedding that captures higher-order topological proximity, meaning an expert’s position in the wider web of past collaborations and answered questions. The authors use a two-layer graph convolutional network, a depth that prior work has shown is sufficient to capture these higher-order relationships without over-smoothing the representations.
The second component is a content encoder that reads the actual text of questions and answers. Rather than relying on exact keyword overlap, which is brittle in the face of vocabulary variation, the model employs a kernel pooling mechanism built on radial basis function kernels. This technique, originally developed for neural ad-hoc ranking in web search, transforms soft word-matching similarities between a new question and an expert’s past answers into a vector of soft term-frequency statistics. Eleven kernels are used: one narrow kernel tuned to capture exact matches, and ten evenly spaced kernels that act as soft bins for graded degrees of similarity. The result is a rich, differentiable signal of how well the language of a new question aligns with the demonstrated expertise of a candidate.
What makes the architecture genuinely novel is how these signals are fused and trained. The outputs of the structure encoder and the content encoder are passed into a multi-layer perceptron ranker that produces final relevance scores, and the entire pipeline, from word embeddings through graph convolutions to the ranking layer, is optimized jointly with gradient descent using the Adam optimizer. Gradients flow backward through the kernel pooling layer into the translation matrix of word similarities and ultimately into the word embeddings themselves, meaning the model learns not just how to rank experts but how to represent language in a way that serves expert discovery. This end-to-end discipline stands in contrast to heuristic team formation methods, which involve iterative combinatorial search and cannot benefit from representation learning at all.
The empirical case for the method rests on experiments across four real-world datasets drawn from Stack Exchange communities: android, history, dba and physics. For individual expert finding, the model outperformed the strongest baselines by at least 4.4 percent in NDCG and 6.7 percent in MAP, two standard ranking metrics that capture, respectively, how closely a predicted ranking matches the ideal ordering of experts and how robustly relevant experts appear across the ranking list. In help-hurt analyses across individual test questions, the proposed method beat the best baseline on 36.4 percent of questions for NDCG and improved MAP over the best baseline by 32.5 percent, indicating that the gains are broad rather than concentrated in a few easy cases.
For team formation, the researchers evaluated three complementary metrics. Skill coverage measures how completely the tags of questions a discovered team has previously answered overlap with the tags of a new question, capturing whether the team collectively possesses the background knowledge required. Collaboration level counts how often team members have answered the same questions in the past, normalized across all pairs, and the gold standard match score measures the fraction of a discovered team that overlaps with the experts who actually answered the question in reality. On skill coverage, the new model achieved at least 4.7 percent broader coverage than collaborative expert finding baselines, and it beat the strongest team formation baselines on 15.2 to 18.4 percent of test questions while losing on only 3.5 to 7.4 percent.
Perhaps the most revealing findings concern trade-offs. The strongest baseline on collaboration level, a method called EnC, tends to retrieve high-degree nodes in the network, users who have answered many questions and therefore share many neighbors, which inflates collaboration scores. But the analysis shows that high collaboration does not guarantee skill coverage: teams formed by EnC had on average 5.91 percent lower skill coverage than those discovered by the new model, while the new model’s teams carried on average 1.7 times higher collaboration than the best team formation baseline. On the gold standard match metric, the new method retrieved actual answerers roughly twice as often as the best baseline and achieved 4.18 times the score of the second-best team formation method, suggesting its teams resemble the real collaborative groups that form organically on these platforms.
Scalability, often the Achilles heel of graph neural approaches, receives careful treatment. The authors show theoretically that the graph convolutional layers scale linearly in the number of edges times the embedding dimension, while the ranker scales with the number of nodes times the squared embedding dimension, a complexity profile comparable to existing neural baselines and free of the iterative search costs of heuristic methods. Empirically, throughput declines as datasets grow, as expected, but even on the largest physics dataset the model processed more than half a sample per second during training and more than one sample per second during inference, speeds that the authors argue are compatible with batch processing and even near real-time expert recommendation in latency-sensitive settings.
Beyond the immediate application to question answering platforms, the work signals a broader shift in how machine learning treats collective expertise. Ablation results show that combining topological and textual signals yields on average 4.4 percent and 8.61 percent better NDCG and MAP than topology alone, although structural data proved the more powerful single signal, improving performance by 16.44 and 19.7 percent over text-only variants. The lesson is that expertise lives simultaneously in what people write and in where they sit within a network of past collaborations, and that models capable of learning both jointly, rather than stitching them together after the fact, are better positioned to find not just who knows the answer, but who can answer it together.
Subject of Research: Joint representation learning with graph neural networks and kernel pooling for discovering individual experts and collaborative expert teams in community question answering
Article Title: Joint Representation Learning for Expert Discovery
Article References: Etemadi, R., Zihayat, M., Feng, K., Adelman, J., Zarrinkalam, F., & Bagheri, E. (2026). Joint Representation Learning for Expert Discovery. Machine Learning, 115(10), Article 234. https://doi.org/10.1007/s10994-026-07166-z
Image Credits: AI Generated
DOI: 10.1007/s10994-026-07166-z
Keywords: expert finding, community question answering, graph neural networks, kernel pooling, team formation, representation learning, question routing, skill coverage, Stack Exchange, machine learning, neural ranking, collaborative expertise
Cite Scienmag News
Denise Maddox. (September 30, 2026). AI Learns to Build Expert Teams for Online Question Answering. Scienmag. https://scienmag.com/ai-learns-to-build-expert-teams-for-online-question-answering/
Denise Maddox. "AI Learns to Build Expert Teams for Online Question Answering." Scienmag, 30 September 2026, https://scienmag.com/ai-learns-to-build-expert-teams-for-online-question-answering/. Accessed 30 September 2026.
Denise Maddox. "AI Learns to Build Expert Teams for Online Question Answering." Scienmag. September 30, 2026. https://scienmag.com/ai-learns-to-build-expert-teams-for-online-question-answering/

