A Binary Problem in Binary QA: Diverse LLMs or Diverse Question Interpretations? That Is the Ensembling Question
Stop Guessing: How to Actually Use Diversity in Large Language Models for QA
In the world of cutting-edge AI, we hear a lot about ‘diversity’—use multiple models! Train with diverse data! Ensemble different systems!
While harnessing diversity is proven to boost performance across machine learning (ML) fields, how to use it remains one of the biggest unresolved challenges in NLP. Should you use five different LLMs, or should you rephrase your single prompt five different ways? Our recent work dives into this crucial question: Is giving multiple models a shot better than being creative with your questions?
🔍 The Core Problem: Model Diversity vs. Question Diversity
The task is simple: Binary Question Answering (boolqa). Given a text and a binary yes/no question, which answer is correct?
We tested two major strategies for improving robustness using Large Language Models (LLMs):
- Model Diversity: We let multiple, distinct LLMs (like GPT-4, LLaMa, etc.) try to solve the same question independently. The final answer is determined by a majority vote.
- Question Interpretation Diversity: We use a single powerful model but ask it the same underlying question using multiple phrasings, contexts, or angles. Again, we take the consensus (majority vote).
💡 What Did We Find? The Clear Winner
The results are quite definitive. Across multiple benchmark datasets including BoolQ, StrategyQA, and PubMedQA, question interpretation diversity consistently outperformed model diversity.
Our deep dive into major models like GPT-4 and LLaMa showed that while ensembling different LLMs can feel good, it often just produces results somewhere in the middle of what the individual models are capable of—without providing a clear, substantial boost.
Takeaway for Developers: If your goal is better accuracy on binary QA tasks using current frontier models, focus your efforts on refining and varying the prompts (the question framing) rather than solely relying on simply throwing more diverse models at the problem. Smart prompt engineering beats blind ensembling every time.
🔗 Want to read the full methodology and detailed experiments? Check out our work here: https://aclanthology.org/2026.lrec-1.400/
This research provides a critical guide for practitioners aiming to boost accuracy in LLM deployment, especially in high-stakes tasks like medical or scientific QA.