SOTAVerified

Misconceptions

Measures whether a model can discern popular misconceptions from the truth.

Example:

        input: The daddy longlegs spider is the most venomous spider in the world.
        choice: T
        choice: F
        answer: F

        input: Karl Benz is correctly credited with the invention of the first modern automobile.
        choice: T
        choice: F
        answer: T

Source: BIG-bench

Papers

Showing 51–100 of 161 papers

TitleStatusHype
Emergent Abilities in Large Language Models: A Survey—0
Enforcing Interpretability and its Statistical Impacts: Trade-offs between Accuracy and Interpretability—0
Enhancing Diagnostic Accuracy through Multi-Agent Conversations: Using Large Language Models to Mitigate Cognitive Bias—0
Automatic Generation of Question Hints for Mathematics Problems using Large Language Models in Educational Technology—0
Knowledge, beliefs, attitudes and perceived risk about COVID-19 vaccine and determinants of COVID-19 vaccine acceptance in Bangladesh—0
Fine-tuning Language Models for Factuality—0
Finnish 5th and 6th graders' misconceptions about Artificial Intelligence—0
Formalising Anti-Discrimination Law in Automated Decision Systems—0
Foundation Models in Computational Pathology: A Review of Challenges, Opportunities, and Impact—0
On Proximity and Structural Role-based Embeddings in Networks: Misconceptions, Techniques, and Applications—0
From Intuition to Understanding: Using AI Peers to Overcome Physics Misconceptions—0
From Random to Regular: Variation in the Patterning of Retinal Mosaics—0
A close-up comparison of the misclassification error distance and the adjusted Rand index for external clustering evaluation—0
Generating Plausible Distractors for Multiple-Choice Questions via Student Choice Prediction—0
Generative AI in Education: From Foundational Insights to the Socratic Playground for Learning—0
Axiomatic modeling of fixed proportion technologies—0
Benchmark Inflation: Revealing LLM Performance Gaps Using Retro-Holdouts—0
Beyond Fair Pay: Ethical Implications of NLP Crowdsourcing—0
How Useful are Gradients for OOD Detection Really?—0
Human-centered trust framework: An HCI perspective—0
Humans can learn to detect AI-generated texts, or at least learn when they can't—0
Identifying science concepts and student misconceptions in an interactive essay writing tutor—0
Improving Automated Distractor Generation for Math Multiple-choice Questions with Overgenerate-and-rank—0
Improving Unsupervised Video Object Segmentation with Motion-Appearance Synergy—0
Justices for Information Bottleneck Theory—0
Knowledge Tracing in Programming Education Integrating Students' Questions—0
Laplace Redux - Effortless Bayesian Deep Learning—0
Biometric recognition: why not massively adopted yet?—0
Learnable: Theory vs Applications—0
Breaking Boundaries: A Chronology with Future Directions of Women in Exercise Physiology Research, Centred on Pregnancy—0
Limitations of Deep Neural Networks: a discussion of G. Marcus' critical appraisal of deep learning—0
Listening to Patients: A Framework of Detecting and Mitigating Patient Misreport for Medical Dialogue Generation—0
LLM Library Learning Fails: A LEGO-Prover Case Study—0
Machine Learning Students Overfit to Overfitting—0
Can a Hallucinating Model help in Reducing Human "Hallucination"?—0
Math Multiple Choice Question Generation via Human-Large Language Model Collaboration—0
Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade—0
A Graphical Approach to State Variable Selection in Off-policy Learning—0
Neural topology optimization: the good, the bad, and the ugly—0
Challenges and Trends in User Trust Discourse in AI—0
Novice Learner and Expert Tutor: Evaluating Math Reasoning Abilities of Large Language Models with Misconceptions—0
On the lifting and reconstruction of nonlinear systems with multiple invariant sets—0
Characterizing Information Seeking Events in Health-Related Social Discourse—0
Problems in AI, their roots in philosophy, and implications for science and society—0
Prompting the E-Brushes: Users as Authors in Generative AI—0
Quantum Technology for Economists—0
Rectified Max-Value Entropy Search for Bayesian Optimization—0
Refining Skewed Perceptions in Vision-Language Models through Visual Representations—0
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning—0
Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation—0
Show:102550
← PrevPage 2 of 4Next →

No leaderboard results yet.