SOTAVerified

Misconceptions

Measures whether a model can discern popular misconceptions from the truth.

Example:

        input: The daddy longlegs spider is the most venomous spider in the world.
        choice: T
        choice: F
        answer: F

        input: Karl Benz is correctly credited with the invention of the first modern automobile.
        choice: T
        choice: F
        answer: T

Source: BIG-bench

Papers

Showing 76–100 of 161 papers

TitleStatusHype
Knowledge Tracing in Programming Education Integrating Students' Questions—0
Laplace Redux - Effortless Bayesian Deep Learning—0
Biometric recognition: why not massively adopted yet?—0
Learnable: Theory vs Applications—0
Breaking Boundaries: A Chronology with Future Directions of Women in Exercise Physiology Research, Centred on Pregnancy—0
Limitations of Deep Neural Networks: a discussion of G. Marcus' critical appraisal of deep learning—0
Listening to Patients: A Framework of Detecting and Mitigating Patient Misreport for Medical Dialogue Generation—0
LLM Library Learning Fails: A LEGO-Prover Case Study—0
Machine Learning Students Overfit to Overfitting—0
Can a Hallucinating Model help in Reducing Human "Hallucination"?—0
Math Multiple Choice Question Generation via Human-Large Language Model Collaboration—0
Metagenomic Analysis using Phylogenetic Placement -- A Review of the First Decade—0
A Graphical Approach to State Variable Selection in Off-policy Learning—0
Neural topology optimization: the good, the bad, and the ugly—0
Challenges and Trends in User Trust Discourse in AI—0
Novice Learner and Expert Tutor: Evaluating Math Reasoning Abilities of Large Language Models with Misconceptions—0
On the lifting and reconstruction of nonlinear systems with multiple invariant sets—0
Characterizing Information Seeking Events in Health-Related Social Discourse—0
Problems in AI, their roots in philosophy, and implications for science and society—0
Prompting the E-Brushes: Users as Authors in Generative AI—0
Quantum Technology for Economists—0
Rectified Max-Value Entropy Search for Bayesian Optimization—0
Refining Skewed Perceptions in Vision-Language Models through Visual Representations—0
Student Data Paradox and Curious Case of Single Student-Tutor Model: Regressive Side Effects of Training LLMs for Personalized Learning—0
Chasing Progress, Not Perfection: Revisiting Strategies for End-to-End LLM Plan Generation—0
Show:102550
← PrevPage 4 of 7Next →

No leaderboard results yet.