SOTAVerified

Chatbot

Chatbot or conversational AI is a language model designed and implemented to have conversations with humans.

Source: Open Data Chatbot

Image source

Papers

Showing 351–375 of 971 papers

TitleStatusHype
Impact of Decoding Methods on Human Alignment of Conversational LLMs—0
Interactive Learning in Computer Science Education Supported by a Discord ChatbotCode0
Enhancing Model Performance: Another Approach to Vision-Language Instruction Tuning—0
Trust or Escalate: LLM Judges with Provable Guarantees for Human Agreement—0
WildHallucinations: Evaluating Long-form Factuality in LLMs with Real-World Entity Queries—0
MoRSE: Bridging the Gap in Cybersecurity Expertise with Retrieval Augmented Generation—0
Impacts of Anthropomorphizing Large Language Models in Learning Environments—0
Chatbot-Based Ontology Interaction Using Large Language Models and Domain-Specific Standards—0
Unipa-GPT: Large Language Models for university-oriented QA in ItalianCode0
Improving Engagement and Efficacy of mHealth Micro-Interventions for Stress Coping: an In-The-Wild Study—0
Trust No Bot: Discovering Personal Disclosures in Human-LLM Conversations in the WildCode0
zIA: a GenAI-powered local auntie assists tourists in Italy—0
Arena Learning: Build Data Flywheel for LLMs Post-training via Simulated Chatbot Arena—0
A Chatbot for Asylum-Seeking Migrants in EuropeCode0
SoupLM: Model Integration in Large Language and Multi-Modal Models—0
Analyzing Large language models chatbots: An experimental approach using a probability test—0
Empirical Study of Symmetrical Reasoning in Conversational Chatbots—0
Towards Optimizing and Evaluating a Retrieval Augmented QA Chatbot using LLMs with Human in the Loop—0
Stephanie: Step-by-Step Dialogues for Mimicking Human Interactions in Social Conversations—0
Zero-shot Persuasive Chatbots with LLM-Generated Strategies and Information Retrieval—0
On the Benchmarking of LLMs for Open-Domain Dialogue Evaluation—0
Improving Conversational Abilities of Quantized Large Language Models via Direct Preference Alignment—0
Model-Enhanced LLM-Driven VUI Testing of VPA Apps—0
Lightweight Large Language Model for Medication Enquiry: Med-Pal—0
Self-Cognition in Large Language Models: An Exploratory Study—0
Show:102550
← PrevPage 15 of 39Next →

Benchmark Results

#ModelMetricClaimedVerifiedStatus
1Yi 34B ChatAverage win rate27.2—Unverified