← Back to Archive

Digest for 2026-10-03

🐦 Share on X 💼 Share on LinkedIn 📘 Share on Facebook

A Comprehensive Evaluation of GenAI-based Items in a Medical Examination

By Yanlin Jiang, Marcus Walker, Andrew Dallas, Aquia Richburg, Nikole Gregg and Brittany Corrigan in Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress • ACL Anthology • Importance: 85/100
Hero Image for acl_2026.aimecon-wip.35

🤖 Passing the Boards: Can GenAI Create Exam-Quality Medical Questions?

The revolution in Generative AI is reshaping nearly every industry, but how do we trust it when stakes are highest—like professional medical examinations? MedEd experts and researchers have been asking this critical question. The answer might just be better than we thought.

The Big Problem: Traditionally, creating high-quality medical exam items requires Subject Matter Experts (SMEs) spending massive amounts of time on manual writing, review, and psychometric validation. This process is slow, expensive, and resource-intensive.

What We Tested: Researchers from the AIME-Con 2026 evaluated Generative AI’s performance by comparing items created by large language models (LLMs) against gold-standard questions designed by seasoned medical experts.

🔬 The Findings: Equally Sharp, Different Speed. The study analyzed both the content of the generated items and the corresponding response data from a high-stakes setting. They found that GenAI-based medical assessment items showed content quality and psychometric characteristics comparable to those developed by human SMEs.

This is huge. It doesn’t just mean AI can write questions; it means they are rigorous enough, reliable enough, and comprehensive enough for critical professional assessments. This finding empirically supports the robust potential of GenAI to revolutionize how we design educational and certification programs across various high-stakes fields (medicine, law, etc.).

🚀 Why Does This Matter for Education & Assessment?

  1. Scalability: Institutions can now generate vast quantities of high-quality assessment material quickly, reducing the bottleneck associated with human expert availability.
  2. Efficiency: It democratizes content creation, allowing smaller or resource-limited programs to maintain top-tier examination standards.
  3. Focus Shift: Human experts can shift their time from tedious item writing to higher-level tasks like validating complex curricula and designing assessment frameworks.

The Takeaway: GenAI isn’t just a novelty feature; it’s a reliable, professional tool poised to fundamentally change the backbone of academic and professional evaluation.

Check out the full findings here: A Comprehensive Evaluation of GenAI-based Items in a Medical Examination

#AIinEducation #EdTech #MedicalAI #GenerativeAI #AssessmentTechnology

A Feasibility Study on Retrieval Practice Using Large Language Models

By Marcus Leong, Jennifer Rose, Lisa Dierker and Antonio Laverghetta Jr. in Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers • ACL Anthology • Importance: 75/100
Hero Image for acl_2026.aimecon-main.21

Is AI Ready to Grade Your Learning? The Truth About LLMs and Retrieval Practice

As Large Language Models (LLMs) become integrated into education technology, the promise of automated learning tools—from personalized quizzes to instant feedback—is immense. But when it comes to core academic functions like generating high-quality review material, does AI really measure up to human experts?

Our latest study dives deep into a critical educational function: Retrieval Practice. Known academically for boosting memory and retention, this practice traditionally requires massive banks of perfectly curated questions (item banks). Historically, building these was a labor-intensive task for educators.

We conducted a head-to-head comparison: we compared items generated by powerful LLMs against quiz questions meticulously written by human instructors in an introductory psychology course. The goal? To determine if AI-generated quizzes maintain the same psychometric quality as professionally designed ones.

🧠 What We Found (The Hard Truth)

The results suggest a cautionary but important conclusion: while LLMs are phenomenal tools for generating drafts and bulk content, their output exhibited overall weaker psychometric properties compared to human-written items.

This doesn’t mean AI is useless! It means that for high-stakes educational measurements—where accuracy of learning assessment is paramount—human supervision and expert refinement remain critical.

🛠️ The Practical Takeaways for EdTech & Educators

  • LLMs as Assistants, Not Replacements: Think of LLMs as powerful co-pilots. They can generate a massive volume of potential items quickly, saving immense time, but they require a domain expert to vet the questions and ensure academic rigor.
  • The Value of Psychometrics: The study underscores that simply having a question isn’t enough; it must be well-designed to truly measure learning outcomes. This requires expertise beyond current LLM capabilities.
  • Future Directions: Integrating AI workflows needs to include mandatory human validation loops, especially for measurement and assessment tasks.

This research provides a crucial feasibility check on automating one of education’s most powerful techniques. For researchers building next-generation adaptive learning platforms, this paper is a must-read: A Feasibility Study on Retrieval Practice Using Large Language Models.


📚 Read the Full Paper: To explore our detailed methodology and full findings, check out the paper here: A Feasibility Study on Retrieval Practice Using Large Language Models.

A Multi-Year Investigation of Undergraduates’ Generative AI Fairness Appraisals in Higher Education

By Sunday Stein, Victoria Delaney and Mich Manrique in Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers • ACL Anthology • Importance: 75/100
Hero Image for acl_2026.aimecon-main.61

Does GenAI Fairness Education Need an Update? Insights from Students.

Generative AI is rapidly reshaping higher education. But how well are students grasping the ethical nuances of fairness in this technology? Our latest research dives deep into this question, offering a critical look at how undergraduate perceptions of ‘GenAI fairness’ evolve over time within academic curricula.

What We Studied:

In a two-year longitudinal study, we tracked undergraduates’ appraisals of generative AI fairness as it was integrated into their coursework. This wasn’t just a single snapshot; we analyzed the shifts in student understanding and rationale year-over-year. Understanding how students think about bias and fairness—and whether those views are maturing—is crucial for updating educational policy.

Key Findings & Why It Matters:

The study reveals measurable differences in how students perceive AI fairness between academic years. These changes aren’t random; they point to specific areas where educational interventions or institutional policies need adjustment. For example, the gap between what institutions teach about AI ethics and what students feel equipped to judge might be widening—or perhaps narrowing—in ways we need to understand.

Implications for Higher Ed Policy:

  • Curriculum Design: Educators can use these insights to fine-tune courses, moving beyond mere awareness toward deep critical appraisal of AI systems. Instead of just using GenAI tools, students learn how and why those tools might be biased.
  • Student Empowerment: The findings underscore the need for curricula that actively engage students in discussing ethical failures in AI, transforming them from passive users into critical evaluators.
  • Policy Action: For university administration, this suggests that GenAI ethics education must be dynamic and responsive, evolving with both technological speed and student maturity.

We encourage higher education institutions and curriculum developers to review their existing AI ethics modules based on these longitudinal findings. Understanding the trajectory of understanding is key to building genuinely fair graduates.

Want to read the full study details?

A Preliminary Semantic Screen for IRT Local Dependence

By Qian Shen in Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Works in Progress • ACL Anthology • Importance: 75/100
Hero Image for acl_2026.aimecon-wip.2

Detecting Flawed Personality Tests: A Semantic Screen for IRT

Are psychometric tests accurately measuring what they’re supposed to? When we use Item Response Theory (IRT) to model personality or ability, one of the biggest risks is local dependence. This means two items in your test might correlate too closely, giving you a skewed understanding of an individual’s true score. It’s like grading an essay using two questions that essentially ask the same thing—you don’t get enough signal.

Our preliminary work tackles this challenge head-on: Can we leverage modern Natural Language Processing (NLP) to detect potential item dependence before complex psychometric modeling begins?

🧠 How We Applied Semantics to Psychometrics

Traditional diagnostics for local dependence often rely on statistical metrics (like covariance), which are great, but they can be computationally heavy or miss nuances rooted in language. Our approach proposes a semantic screening layer. Instead of just looking at the correlation coefficient between items, we use advanced sentence embeddings to measure how semantically similar item pairs are.

We applied this method to a standard 50-item personality inventory dataset. The results show a preliminary link: highly similar semantic embedding scores often flag item pairs that are statistically more likely to exhibit residual local dependence in the IRT model.

The takeaway? Before running millions of data points through your complex psychometric pipeline, you can quickly filter out redundant or overlapping items using pure NLP metrics. This significantly improves both the speed and validity of your underlying measurements.

📈 What Does This Mean for Researchers?

  1. Increased Validity: By identifying dependent items early, researchers ensure their core measurement model is built on truly independent variables.
  2. Efficiency Boost: Semantic screening acts as a fast pre-processor, saving computational time and refining the item bank from the start.
  3. Multidimensional Models: This technique is particularly valuable for complex multidimensional IRT models where every piece of redundant data can introduce noise or bias.

This research offers a powerful conceptual bridge between cutting-edge NLP (like transformer embeddings) and established psychometrics, opening new pathways for creating more robust and linguistically informed assessment tools. For more details on our findings, check out the full working paper: A preliminary semantic screen for IRT Local Dependence.

^(Disclaimer: This is a preliminary work-in-progress analysis, suggesting further testing across diverse datasets and samples.)

A Validity Threat Framework for Measuring Student Generative AI Use

By Olukayode Emmanuel Apata, Yetunde Omoyiwola Fawehinmi, Daniel Olutola Oyeniran, Glory Onize Saidu, Segun Timothy Ajose, Naphtali Onalo and Oluwasegun Matthew Amoniyan in Proceedings of the Artificial Intelligence in Measurement and Education Conference (AIME-Con): Full Papers • ACL Anthology • Importance: 75/100
Hero Image for acl_2026.aimecon-main.25

🤖 AI in Education: A New Framework for Measuring Student GenAI Use

As Generative AI tools like ChatGPT and Midjourney become standard classroom aids, measuring how students use them—and whether that usage undermines learning validity—is a massive challenge. Traditional assessment methods are struggling to keep up with the pace of technological change.

This paper introduces a comprehensive Validity Threat Framework specifically designed for higher education. It provides academic and educators with a critical lens to evaluate assessments in an AI-integrated world. Instead of just saying ‘AI cheating,’ it dissects why and where measurement efforts might fail.

🔍 What is the Validity Threat Framework?

The core value lies in its breadth. It doesn’t treat GenAI usage as a single problem, but addresses multiple dimensions where validity could be compromised:

  • Construct Definition: How are we even defining ‘AI use’? Is it simply writing prompts, or does it include generating ideas and structuring arguments? The framework helps refine this boundary.
  • Response Processes & Self-Report Bias: It warns that student self-reports of AI usage can be unreliable (students may overreport or underreport).
  • Policy Context & Fairness: Assessments must account for existing institutional policies and ensure the measurements are fair across different demographics and learning styles.
  • Score Interpretation: Even if we measure use, how do we interpret that score? Does a high usage score mean low performance, or simply heavy resource utilization?

💡 Practical Takeaways for Educators & Policy Makers

The authors move beyond theory by providing concrete guidance. This isn’t just an academic critique; it’s a roadmap for action in AI-integrated education.

  1. Redesign Assessments: Shift away from easily AI-generated tasks (like simple essays) toward process-based assignments, personalized problem sets, and oral defenses that require unique student synthesis.
  2. Develop Clear Policies: Institutions need transparent guidelines on what constitutes acceptable use versus academic misconduct when using GenAI.
  3. Improve Pedagogy: Educators must integrate AI usage into the learning process itself—treating it as a skill to be taught, not just a threat to be banned.

This conceptual work is essential reading for curriculum designers, educational psychometricians, and any professional grappling with assessment validity in the age of ChatGPT. Dive deeper into the full discussion here: Validity Threat Framework for GenAI Use


#EdTech #HigherEducation #GenAI #AssessmentDesign #AIinLearning

Explore Recent Digests