Difference between revisions of "Benchmarks"

From
Jump to: navigation, search
m
m (Status: Passed, and Reinterpreted)
 
(5 intermediate revisions by the same user not shown)
Line 1: Line 1:
 +
__NOTOC__
 
{{#seo:
 
{{#seo:
 
|title=PRIMO.ai
 
|title=PRIMO.ai
 
|titlemode=append
 
|titlemode=append
|keywords=ChatGPT, artificial, intelligence, machine, learning, GPT-4, GPT-5, NLP, NLG, NLC, NLU, models, data, singularity, moonshot, Sentience, AGI, Emergence, Moonshot, Explainable, TensorFlow, Google, Nvidia, Microsoft, Azure, Amazon, AWS, Hugging Face, OpenAI, Tensorflow, OpenAI, Google, Nvidia, Microsoft, Azure, Amazon, AWS, Meta, LLM, metaverse, assistants, agents, digital twin, IoT, Transhumanism, Immersive Reality, Generative AI, Conversational AI, Perplexity, Bing, You, Bard, Ernie, prompt Engineering LangChain, Video/Image, Vision, End-to-End Speech, Synthesize Speech, Speech Recognition, Stanford, MIT |description=Helpful resources for your journey with artificial intelligence; videos, articles, techniques, courses, profiles, and tools 
+
|keywords=benchmarks, evaluation, Turing test, imitation game, Chinese Room, LLM evaluation, GLUE, SuperGLUE, SQuAD, MMLU, HELM, BIG-bench, ARC, WinoGrande, HumanEval, GSM8K, GPQA, ARC-AGI, SWE-bench, FrontierMath, Humanity's Last Exam, agentic evaluation, contamination, leaderboard, Procgen, MLPerf, CAPTCHA, reCAPTCHA, LegalBench, clinical NLP, artificial intelligence, machine learning, models, data
 +
|description=Named tests, datasets, and leaderboards used to measure AI systems — from the Turing test through language understanding benchmarks, agentic evaluation, hardware suites, and the limits of each
 +
}}
  
<!-- Google tag (gtag.js) -->
+
[https://www.youtube.com/results?search_query=ai+~Benchmark YouTube] [https://www.quora.com/search?q=ai%20~Benchmark ... Quora] [https://www.google.com/search?q=ai+~Benchmark ...Google search] [https://news.google.com/search?q=ai+~Benchmark ...Google News] [https://www.bing.com/news/search?q=ai+~Benchmark&qft=interval%3d%228%22 ...Bing News]
<script async src="https://www.googletagmanager.com/gtag/js?id=G-4GCWLBVJ7T"></script>
 
<script>
 
  window.dataLayer = window.dataLayer || [];
 
  function gtag(){dataLayer.push(arguments);}
 
  gtag('js', new Date());
 
 
 
  gtag('config', 'G-4GCWLBVJ7T');
 
</script>
 
}}
 
[https://www.youtube.com/results?search_query=ai+~Benchmark YouTube]
 
[https://www.quora.com/search?q=ai%20~Benchmark ... Quora]
 
[https://www.google.com/search?q=ai+~Benchmark ...Google search]
 
[https://news.google.com/search?q=ai+~Benchmark ...Google News]
 
[https://www.bing.com/news/search?q=ai+~Benchmark&qft=interval%3d%228%22 ...Bing News]
 
  
* [[Data Science]] ... [[Data Governance|Governance]] ... [[Data Preprocessing|Preprocessing]] ... [[Feature Exploration/Learning|Exploration]] ... [[Data Interoperability|Interoperability]] ... [[Algorithm Administration#Master Data Management (MDM)|Master Data Management (MDM)]] ... [[Bias and Variances]] ... [[Benchmarks]] ... [[Datasets]]  
+
* [[Data Science]] ... [[Data Governance|Governance]] ... [[Data Preprocessing|Preprocessing]] ... [[Feature Exploration/Learning|Exploration]] ... [[Data Interoperability|Interoperability]] ... [[Algorithm Administration#Master Data Management (MDM)|Master Data Management (MDM)]] ... [[Bias and Variances]] ... Benchmarks ... [[Datasets]]
* [[Large Language Model (LLM)]] ... [[Natural Language Processing (NLP)]] ... [[Natural Language Generation (NLG)|Generation]] ... [[Natural Language Classification (NLC)|Classification]] ... [[Natural Language Processing (NLP)#Natural Language Understanding (NLU)|Understanding]] ... [[Language Translation|Translation]] ... [[Natural Language Tools & Services|Tools & Services]]
+
* [[Large Language Model (LLM)]] ... [[Natural Language Processing (NLP)]] ... [[Natural Language Generation (NLG)|Generation]] ... [[Natural Language Classification (NLC)|Classification]] ... [[Natural Language Processing (NLP)#Natural Language Understanding (NLU)|Understanding]] ... [[Language Translation|Translation]] ... [[Natural Language Tools & Services|Tools & Services]]
 
* [[Risk, Compliance and Regulation]] ... [[Ethics]] ... [[Privacy]] ... [[Law]] ... [[AI Governance]] ... [[AI Verification and Validation]]
 
* [[Risk, Compliance and Regulation]] ... [[Ethics]] ... [[Privacy]] ... [[Law]] ... [[AI Governance]] ... [[AI Verification and Validation]]
 
* [[Case Studies]]
 
* [[Case Studies]]
 
** [[Gaming]]
 
** [[Gaming]]
 
** [[Algorithm Administration#Model Monitoring|Model Monitoring]]
 
** [[Algorithm Administration#Model Monitoring|Model Monitoring]]
* [[ALFRED]] ... Action Learning From Realistic Environments and Directives
+
* [[ALFRED]] ... Action Learning From Realistic Environments and Directives
 
* [[Algorithm Administration]]
 
* [[Algorithm Administration]]
* [[Data Quality]] ...[[AI Verification and Validation|validity]], [[Evaluation - Measures#Accuracy|accuracy]], [[Data Quality#Data Cleaning|cleaning]], [[Data Quality#Data Completeness|completeness]], [[Data Quality#Data Consistency|consistency]], [[Data Quality#Data Encoding|encoding]], [[Data Quality#Zero Padding|padding]], [[Data Quality#Data Augmentation, Data Labeling, and Auto-Tagging|augmentation, labeling, auto-tagging]], [[Data Quality#Batch Norm(alization) & Standardization| normalization, standardization]], and [[Data Quality#Imbalanced Data|imbalanced data]]
+
* [[Data Quality]] ...[[AI Verification and Validation|validity]], [[Evaluation - Measures#Accuracy|accuracy]], [[Data Quality#Data Cleaning|cleaning]], [[Data Quality#Data Completeness|completeness]], [[Data Quality#Data Consistency|consistency]], [[Data Quality#Data Encoding|encoding]], [[Data Quality#Zero Padding|padding]], [[Data Quality#Data Augmentation, Data Labeling, and Auto-Tagging|augmentation, labeling, auto-tagging]], [[Data Quality#Batch Norm(alization) & Standardization|normalization, standardization]], and [[Data Quality#Imbalanced Data|imbalanced data]]
* [[Natural Language Processing (NLP)#Managed Vocabularies |Managed Vocabularies]]
+
* [[Natural Language Processing (NLP)#Managed Vocabularies|Managed Vocabularies]]
 
* [[Excel]] ... [[LangChain#Documents|Documents]] ... [[Database|Database; Vector & Relational]] ... [[Graph]] ... [[LlamaIndex]]
 
* [[Excel]] ... [[LangChain#Documents|Documents]] ... [[Database|Database; Vector & Relational]] ... [[Graph]] ... [[LlamaIndex]]
 
* [[Analytics]] ... [[Visualization]] ... [[Graphical Tools for Modeling AI Components|Graphical Tools]] ... [[Diagrams for Business Analysis|Diagrams]] & [[Generative AI for Business Analysis|Business Analysis]] ... [[Requirements Management|Requirements]] ... [[Loop]] ... [[Bayes]] ... [[Network Pattern]]
 
* [[Analytics]] ... [[Visualization]] ... [[Graphical Tools for Modeling AI Components|Graphical Tools]] ... [[Diagrams for Business Analysis|Diagrams]] & [[Generative AI for Business Analysis|Business Analysis]] ... [[Requirements Management|Requirements]] ... [[Loop]] ... [[Bayes]] ... [[Network Pattern]]
Line 44: Line 33:
 
* [https://www.datasciencecentral.com/group/resources/forum/topics/benchmarking-20-machine-learning-models-accuracy-and-speed Benchmarking 20 Machine Learning Models Accuracy and Speed | Marc Borowczak - Data Science Central]
 
* [https://www.datasciencecentral.com/group/resources/forum/topics/benchmarking-20-machine-learning-models-accuracy-and-speed Benchmarking 20 Machine Learning Models Accuracy and Speed | Marc Borowczak - Data Science Central]
 
* [https://www.sciencedirect.com/science/article/pii/S1532046418300716 Benchmarking deep learning models on large healthcare datasets | S. Purushotham, C. Meng, Z. Chea, and Y. Liu]
 
* [https://www.sciencedirect.com/science/article/pii/S1532046418300716 Benchmarking deep learning models on large healthcare datasets | S. Purushotham, C. Meng, Z. Chea, and Y. Liu]
* [https://spectrum.ieee.org/ai-supercomputer#toggle-gdpr Supercomputers Flex Their AI Muscles New benchmarks reveal science-task speedups | Sammuel K. Moore - IEEE Spectrum]
+
* [https://spectrum.ieee.org/ai-supercomputer#toggle-gdpr Supercomputers Flex Their AI Muscles: New benchmarks reveal science-task speedups | Samuel K. Moore - IEEE Spectrum]
* [https://towardsdatascience.com/the-olympics-of-ai-benchmarking-machine-learning-systems-c4b2051fbd2b The Olympics of AI: Benchmarking Machine Learning Systems | Matthew Stewart - Towards Data Science - Medium] ... How do benchmarks birth breakthroughs?
+
* [https://towardsdatascience.com/the-olympics-of-ai-benchmarking-machine-learning-systems-c4b2051fbd2b The Olympics of AI: Benchmarking Machine Learning Systems | Matthew Stewart - Towards Data Science] ... How do benchmarks birth breakthroughs?
 +
 
 +
Three pages on this wiki cover overlapping ground. The division of labor:
  
 +
* '''Benchmarks''' (this page) — named tests, datasets, suites, and leaderboards. GLUE, MMLU, MLPerf, the Turing test.
 +
* '''[[Evaluation]]''' — the process and methodology of assessing a system or a project.
 +
* '''[[Evaluation - Measures]]''' — the metrics themselves: [[Evaluation - Measures#Accuracy|accuracy]], [[Evaluation - Measures#Precision & Recall (Sensitivity)|precision and recall]], [[Evaluation - Measures#Specificity|specificity]].
  
<hr><center><b><i>
+
A recurring caution applies to everything below. '''Benchmarks saturate.''' The history of the field is a sequence of tests that fell and were then dismissed as never having measured intelligence — checkers, chess, Jeopardy!, Go, protein structure prediction, competition mathematics, and most recently the Turing test itself. A high score is evidence about a test, not a verdict on a system. The constraints that make a benchmark meaningful — contamination resistance, construct validity, independent administration — are discussed on [[AI Verification and Validation]].
 +
<br><br>
 +
<hr><center>
  
You can’t improve what you don’t measure.</i></b> — Peter Drucker
+
<b><i>You can’t improve what you don’t measure.</i></b> — Peter Drucker
  
 
</center><hr>
 
</center><hr>
 +
<br>
  
 +
= The Turing Test and Related Thought Experiments =
  
 +
[https://www.youtube.com/results?search_query=ai+~Consciousness+test YouTube] [https://www.quora.com/search?q=ai%20~Consciousness%20test ... Quora] [https://www.google.com/search?q=ai+~Consciousness+test ...Google search] [https://news.google.com/search?q=ai+~Consciousness+test ...Google News] [https://www.bing.com/news/search?q=ai+~Consciousness+test&qft=interval%3d%228%22 ...Bing News]
  
= <span id="AI Consciousness Testing"></span>AI Consciousness Testing =
+
* [[Artificial General Intelligence (AGI) to Singularity]] ... [[Inside Out - Curious Optimistic Reasoning|Curious Reasoning]] ... [[Emergence]] ... [[Moonshots]] ... [[Explainable / Interpretable AI|Explainable AI]] ... [[Algorithm Administration#Automated Learning|Automated Learning]]
[https://www.youtube.com/results?search_query=ai+~Consciousness+test YouTube]
+
* [[AI Verification and Validation]] ... on why passing a test of indistinguishability is not the same as verifying a capability
[https://www.quora.com/search?q=ai%20~Consciousness%20test ... Quora]
+
* [[Perspective]] ... [[Context]] ... [[In-Context Learning (ICL)]] ... [[Transfer Learning]] ... [[Out-of-Distribution (OOD) Generalization]]
[https://www.google.com/search?q=ai+~Consciousness+test ...Google search]
+
* [[Causation vs. Correlation]] ... [[Autocorrelation]] ...[[Convolution vs. Cross-Correlation (Autocorrelation)]]
[https://news.google.com/search?q=ai+~Consciousness+test ...Google News]
 
[https://www.bing.com/news/search?q=ai+~Consciousness+test&qft=interval%3d%228%22 ...Bing News]
 
  
* [[Artificial General Intelligence (AGI) to Singularity]] ... [[Inside Out - Curious Optimistic Reasoning| Curious Reasoning]] ... [[Emergence]] ... [[Moonshots]] ... [[Explainable / Interpretable AI|Explainable AI]] ...  [[Algorithm Administration#Automated Learning|Automated Learning]]
+
== Turing Test ==
* [[Artificial Consciousness / Sentience#Theory of Mind (ToM)|Theory of Mind (ToM)]]
 
* [[Perspective]] ... [[Context]] ... [[In-Context Learning (ICL)]] ... [[Transfer Learning]] ... [[Out-of-Distribution (OOD) Generalization]]
 
* [[Causation vs. Correlation]] ... [[Autocorrelation]] ...[[Convolution vs. Cross-Correlation (Autocorrelation)]]
 
  
== <span id="Turing Test"></span>Turing Test ==
 
 
* [https://www.youtube.com/results?search_query=Alan+Turing YouTube]
 
* [https://www.youtube.com/results?search_query=Alan+Turing YouTube]
 
* [[Creatives#Alan Turing|Alan Turing]]
 
* [[Creatives#Alan Turing|Alan Turing]]
  
The Turing test, originally called the imitation game by Alan Turing in 1950, is a test of a machine's ability to exhibit intelligent behavior equivalent to, or indistinguishable from, that of a human. Turing proposed that a human evaluator would judge natural language conversations between a human and a machine designed to generate human-like responses. The evaluator would be aware that one of the two partners in conversation was a machine, and all participants would be separated from one another. The conversation would be limited to a text-only channel, such as a computer keyboard and screen, so the result would not depend on the machine's ability to render words as speech. If the evaluator could not reliably tell the machine from the human, the machine would be said to have passed the test. The test results would not depend on the machine's ability to give correct answers to questions, only on how closely its answers resembled those a human would give. - [https://en.wikipedia.org/wiki/Turing_test Turing Test | Wikipedia]
+
The Turing test, originally called the imitation game by Alan Turing in 1950, is a test of a machine's ability to exhibit intelligent behavior equivalent to, or indistinguishable from, that of a human. A human evaluator judges natural language conversations between a human and a machine designed to generate human-like responses. The evaluator knows that one of the two partners is a machine, and all participants are separated from one another. The conversation is limited to a text-only channel so the result does not depend on the machine's ability to render words as speech. If the evaluator cannot reliably tell the machine from the human, the machine is said to have passed. The result does not depend on whether the machine gives correct answers, only on how closely its answers resemble those a human would give. - [https://en.wikipedia.org/wiki/Turing_test Turing Test | Wikipedia]
 +
 
 +
=== Status: Passed, and Reinterpreted ===
 +
 
 +
The benchmark fell in a peer-reviewed study published in the Proceedings of the National Academy of Sciences by Cameron Jones and Benjamin Bergen of UC San Diego, using the standard three-party format Turing described. [https://www.pnas.org/doi/10.1073/pnas.2524472123 Large language models pass a standard three-party Turing test | PNAS]
  
 +
What the study found:
  
<hr><center>
+
* '''GPT-4.5 given a persona prompt was judged human 73% of the time''' — significantly above chance, and more often than the actual human participants it was paired against.
 +
* '''ELIZA, the 1960s chatbot, scored 23%''' as a manipulation check, confirming that interrogators and the design were sensitive enough to detect a machine and that the result was not produced by random guessing.
 +
* '''Without the persona prompt the same models did not robustly pass.''' In some conditions their pass rates were not significantly better than ELIZA's.
 +
* A third preregistered study used a '''15-minute''' limit to test whether models continue to pass under extended interrogation.
  
<i><b>Today an AI has to dumb down to pass the Turing Test</b></i> - Ray Kurzweil
+
What it does not establish:
  
</center><hr>
+
* '''Not intelligence.''' The authors frame the test as a measure of ''substitutability'' — whether a system can stand in for a person without the difference being noticed.
 +
* '''Not consciousness.''' Indistinguishability is a fact about observers, not about the system observed.
 +
* '''Not an unaided model result.''' Prompting and scaffolding carried much of the outcome, which is exactly the confound described under [[AI Verification and Validation]].
  
 +
The authors' stated concern is social and economic rather than philosophical: systems that pass as human enable "counterfeit people," with consequences for online trust, employment, genuine social engagement, and the perceived value of human interaction.
  
{|<!-- T -->
+
<br>
| valign="top" |
+
<hr><center>
{| class="wikitable" style="width: 550px;"
 
||
 
<youtube>4VROUIAF2Do</youtube>
 
<b>What is a Turing Test? A Brief History of the Turing Test and its Impact
 
</b><br>[https://www.techtarget.com/searchenterpriseai/definition/Turing-test What is a Turing Test]
 
  
Is a computer as smart as a human? Only a Turing Test will tell -- plus its many spin-offs. A Turing Test is a method of determining whether a computer is capable of thinking like a human. Watch to learn what a Turing Test is and how it relates to AI technology.
+
<b><i>Today an AI has to dumb down to pass the Turing Test</i></b> - Ray Kurzweil
|}
 
|<!-- M -->
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
||
 
<youtube>_GCTLciqT0A</youtube>
 
<b>Will [[ChatGPT]] Pass The Turing Test? Let's Find Out!
 
</b><br>I have been testing [[ChatGPT]] for the past few days and it has been nothing short of spectacular. Now the moment of truth is upon is: Will it pass the Turing Test? Find out in this video. Will it exhibit intelligence that will fool humans into thinking it's not a machine? You'd be surprised!
 
  
[[ChatGPT]] says:
+
</center><hr><br>
  
"The Turing Test is a measure of a machine's ability to exhibit intelligent behavior that is indistinguishable from that of a human. It was first proposed by the British mathematician and computer scientist Alan Turing in 1950. The basic idea of the test is that a human evaluator engages in a text-based conversation with both a human and a machine, without knowing which is which. If the evaluator is unable to reliably determine which is the human and which is the machine, then the machine is said to have passed the Turing Test and demonstrated human-like intelligence.
+
<youtube>4VROUIAF2Do</youtube> <youtube>_GCTLciqT0A</youtube>
  
The Turing Test has become an influential concept in the field of artificial intelligence and continues to be an active area of research and development. While some AI systems have been able to fool evaluators into thinking they are human in limited cases, no machine has yet passed the Turing Test in a comprehensive and sustained manner. Nonetheless, the Turing Test remains a useful benchmark for evaluating the progress of AI and a means for stimulating discussion about the nature of human intelligence and the potential for machines to possess similar capabilities."
+
== Chinese Room Thought Experiment ==
|}
 
|}<!-- B -->
 
  
== <span id="Chinese Room Thought Experiment"></span>Chinese Room Thought Experiment ==
 
 
* [[Creatives#John Searle|John Searle]]
 
* [[Creatives#John Searle|John Searle]]
  
The Chinese Room Experiment is a thought experiment proposed by John Searle in 1980 to argue against the claim that a computer can have a mind or be conscious. Searle's argument has been criticized by some philosophers and computer scientists. However, it remains a powerful argument against the claim that computers can have a mind or be conscious.
+
The Chinese Room is a thought experiment, not a test anyone administers. It belongs here as context for what the Turing test does and does not settle; the broader questions it raises about machine minds are treated under [[Emergence]] and [[Artificial General Intelligence (AGI) to Singularity]].
  
In the experiment, Searle imagines himself locked in a room with a set of rules for manipulating Chinese symbols. The rules are written in English, which Searle understands, but the Chinese symbols are meaningless to him. He is given Chinese characters on slips of paper, which he then processes according to the rules. He then produces Chinese characters on slips of paper in response. To an outside observer, it would appear that Searle understands Chinese and is having a conversation with them. However, Searle himself does not understand Chinese at all. He is simply following the rules blindly.
+
Searle proposed it in 1980 to argue against the claim that a computer can have a mind. He imagines himself locked in a room with a set of rules, written in English, for manipulating Chinese symbols that are meaningless to him. He receives Chinese characters, processes them according to the rules, and produces Chinese characters in response. To an outside observer he appears to understand Chinese. He does not; he is following rules blindly.
  
Searle argues that this shows that a computer, which is essentially a machine that follows rules, cannot be said to understand Chinese or to have a mind. The computer may be able to produce intelligent-sounding output, but it does not have the same kind of understanding that a human being has.
+
Key points of the argument:
  
The Chinese Room Experiment has been widely discussed and debated by philosophers and computer scientists. Some have argued that Searle's argument is flawed, while others have agreed with his conclusion. The Chinese Room Experiment is a complex and challenging thought experiment, and there is no easy answer to the question of whether or not it succeeds in its goal. However, it is a thought-provoking experiment that has helped to shape the debate about artificial intelligence and the nature of mind. Here are some of the key points of Searle's argument:
+
* Understanding a language is not just manipulating symbols according to rules. It also requires a grasp of the meaning of the symbols.
 +
* A computer can manipulate symbols, but does not have the same kind of understanding a human being has.
 +
* The experiment shows that a computer cannot be said to understand Chinese even if it produces intelligent-sounding output.
  
* Understanding a language is not just about manipulating symbols according to rules. It also requires having a grasp of the meaning of the symbols.
+
The argument has been widely disputed by both philosophers and computer scientists, and there is no settled answer as to whether it succeeds. It remains useful here for one reason: it separates ''behaving as if'' from ''being'', which is precisely the gap the Turing test cannot close.
* A computer can manipulate symbols, but it does not have the same kind of understanding that a human being has.
 
* The Chinese Room Experiment shows that a computer cannot be said to understand Chinese, even if it can produce intelligent-sounding output.
 
  
<youtube>tBE06SdgzwM</youtube>
+
<youtube>tBE06SdgzwM</youtube> <youtube>rHKwIYsPXLg</youtube>
<youtube>rHKwIYsPXLg</youtube>
 
  
= <span id="Large Language Model (LLM) Evaluation"></span>Large Language Model (LLM) Evaluation =
+
= Large Language Model (LLM) Evaluation =
[https://www.youtube.com/results?search_query=Evaluat+LLM+Large+Language+Model+Harness+framework YouTube]
 
[https://www.quora.com/search?q=Evaluat%20LLM%20Large%20Language%20Model%20Harness%20framework ... Quora]
 
[https://www.google.com/search?q=Evaluat+LLM+Large+Language+Model+Harness+framework ...Google search]
 
[https://news.google.com/search?q=Evaluat+LLM+Large+Language+Model+Harness+framework ...Google News]
 
[https://www.bing.com/news/search?q=Evaluat+LLM+Large+Language+Model+Harness+framework&qft=interval%3d%228%22 ...Bing News]
 
  
* [[Large Language Model (LLM)]]
+
[https://www.youtube.com/results?search_query=Evaluat+LLM+Large+Language+Model+Harness+framework YouTube] [https://www.quora.com/search?q=Evaluat%20LLM%20Large%20Language%20Model%20Harness%20framework ... Quora] [https://www.google.com/search?q=Evaluat+LLM+Large+Language+Model+Harness+framework ...Google search] [https://news.google.com/search?q=Evaluat+LLM+Large+Language+Model+Harness+framework ...Google News] [https://www.bing.com/news/search?q=Evaluat+LLM+Large+Language+Model+Harness+framework&qft=interval%3d%228%22 ...Bing News]
* [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[Bing/Copilot]] | [[Microsoft]] ... [[Gemini]] | [[Google]] ... [[Claude]] | [[Anthropic]] ... [[Perplexity]] ... [[You]] ... [[phind]] ... [[Ernie]] | [[Baidu]]
+
 
* [[Claude]] | [[Anthropic]]
+
* [[Large Language Model (LLM)]] ... [[Foundation Models (FM)]] ... [[Datasets]] ... [[Evaluation]]
 +
* [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[Bing/Copilot]] | [[Microsoft]] ... [[Gemini]] | [[Google]] ... [[Claude]] | [[Anthropic]] ... [[Perplexity]] ... [[You]] ... [[Phind|phind]] ... [[Ernie]] | [[Baidu]] ... [[DeepSeek]]
 
* [[Large Language Model (LLM)#LLM Token / Parameter / Weight|LLM Token / Parameter / Weight]]
 
* [[Large Language Model (LLM)#LLM Token / Parameter / Weight|LLM Token / Parameter / Weight]]
 
* [[In-Context Learning (ICL)]] ... [[Context]] ... [[Causation vs. Correlation]] ... [[Autocorrelation]] ... [[Out-of-Distribution (OOD) Generalization]] ... [[Transfer Learning]]
 
* [[In-Context Learning (ICL)]] ... [[Context]] ... [[Causation vs. Correlation]] ... [[Autocorrelation]] ... [[Out-of-Distribution (OOD) Generalization]] ... [[Transfer Learning]]
 
* [https://crfm.stanford.edu/helm/latest/ Holistic Evaluation of Language Models (HELM) | Stanford] ... a living benchmark that aims to improve the transparency of language models.
 
* [https://crfm.stanford.edu/helm/latest/ Holistic Evaluation of Language Models (HELM) | Stanford] ... a living benchmark that aims to improve the transparency of language models.
 
* [https://www.mosaicml.com/blog/llm-evaluation-for-icl Blazingly Fast LLM Evaluation for In-Context Learning | Jeremy Dohmann - Mosaic]
 
* [https://www.mosaicml.com/blog/llm-evaluation-for-icl Blazingly Fast LLM Evaluation for In-Context Learning | Jeremy Dohmann - Mosaic]
* [https://github.com/openai/evals Evals - GitHub] ... a framework for evaluating LLMs (large language models) or systems built using LLMs as components.
+
* [https://github.com/openai/evals Evals - GitHub] ... a framework for evaluating LLMs or systems built using LLMs as components.
* [https://wandb.ai/wandb_gen/llm-evaluation/reports/Evaluating-Large-Language-Models-LLMs-with-Eleuther-AI--VmlldzoyOTI0MDQ3 Evaluating Large Language Models (LLMs) with Eleuther AI | Bharat Ramanathan - Weights & Biases] ... With a flexible and tokenization-agnostic interface, the lm-eval library provides a single framework for evaluating and reporting auto-regressive language models on various [[Natural Language Processing (NLP)#Natural Language Understanding (NLU)| Natural Language Understanding (NLU)]] tasks. There are currently over 200 evaluation tasks that support the evaluation of models such as GPT-2 ,T5, Gpt-J, Gpt-Neo, Gpt-NeoX, Flan-T5.
+
* [https://wandb.ai/wandb_gen/llm-evaluation/reports/Evaluating-Large-Language-Models-LLMs-with-Eleuther-AI--VmlldzoyOTI0MDQ3 Evaluating Large Language Models (LLMs) with Eleuther AI | Bharat Ramanathan - Weights & Biases] ... the lm-eval library provides a tokenization-agnostic framework for evaluating auto-regressive language models across [[Natural Language Processing (NLP)#Natural Language Understanding (NLU)|Natural Language Understanding (NLU)]] tasks, with over 200 evaluation tasks.
  
Benchmarks for an LLM:
+
== Dimensions to Compare ==
* <b>Ability to add attachments to prompts</b>: attachments, such as images or documents, the use of attachments allows an LLM to incorporate additional information beyond the textual prompt, which can improve its ability to generate accurate and relevant responses. Claude 2: prompts can include attachments
 
* <b>Performance on the bar exam multiple-choice section</b>: The bar exam is a standardized test that is required to practice law in the United States. The multiple-choice section tests knowledge of legal concepts and principles. Claude 2: Scored 76.5%
 
* <b>Performance on the GRE reading and writing exams</b>: a standardized test that is often required for admission to graduate programs. The reading and writing sections test reading comprehension, analytical writing, and critical thinking skills. Claude 2: A score above the 90th percentile indicates that the LLM is highly proficient in these skills.
 
* <b>Performance on the GRE quantitative reasoning exam</b>: This section tests mathematical and analytical skills. Claude 2: A score similar to the median applicant indicates that the LLM has average proficiency in these skills.
 
* <b>Input length limit</b>: the maximum length of the input prompt that an LLM can handle. A token is a sequence of characters that represents a unit of meaning in natural language processing. Claude 2: A limit of 100K tokens per prompt means that the LLM can handle prompts of up to 100,000 tokens in length.
 
* <b>Context window limit</b>: Maximum amount of context that an LLM can consider when generating a response to a prompt. Claude 2: A context window of up to 100K means that the LLM can consider up to 100,000 tokens of context when generating a response.
 
* <b>Code Generation on HumanEval</b>: This test evaluates the LLM's ability to write code that meets certain criteria, such as correctness and efficiency. Claude 2 on Python coding test: 71.2%
 
* <b>GSM8k math problem set</b>: This problem set evaluates the LLM's ability to solve mathematical problems of varying difficulty. Claude 2: 88%
 
  
 +
These are the axes along which language models are usually compared. '''Specific model scores are not listed here''' — they go stale within months and belong on the individual model pages ([[Claude]] | [[Anthropic]], [[ChatGPT]] | [[OpenAI]], [[Gemini]] | [[Google]], and so on).
  
There are several factors that should be considered while evaluating Large Language Models (LLMs). These include:
+
* '''Attachment support''' — whether prompts can include images or documents, allowing the model to incorporate information beyond the text prompt.
* authenticity
+
* '''Input length limit''' — the maximum prompt length the model accepts, measured in tokens.
* speed
+
* '''[[Context]] window''' — the maximum amount of context the model can consider when generating a response.
* grammar
+
* '''Professional examinations''' — bar exam multiple choice, GRE verbal and quantitative, medical licensing. Tests domain knowledge under a standardized rubric; heavily exposed to training-data contamination.
* readability
+
* '''Code generation''' — HumanEval and successors, measuring whether generated code passes hidden tests.
* unbiasedness
+
* '''Mathematical reasoning''' — GSM8K for grade-school word problems, harder sets for competition and research mathematics.
* backtracking  
+
* '''Tool use and long-horizon tasks''' — see [[#Frontier Evaluations|Frontier Evaluations]] below.
* safety
+
 
* responsibility
+
Qualitative factors that resist single-number scoring: authenticity, speed, grammar, readability, unbiasedness, backtracking behavior, safety, responsibility, contextual understanding, and text operations.
* understanding the context
 
* text operations
 
  
 
== A Survey on Evaluation of Large Language Models ==
 
== A Survey on Evaluation of Large Language Models ==
* [https://arxiv.org/pdf/2307.03109.pdf A Survey on Evaluation of Large Language Models | Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, W. Ye, Y. Zhang, Y. Chang,P. Yu, Q. Yang, X. Xie - ARxIV] ... The survey covers seven major categories of LLM trustworthiness:
 
** <b>Reliability</b>: LLMs should be able to consistently generate accurate and truthful outputs, even when presented with new or challenging inputs.
 
** <b>Safety</b>: LLMs should not generate outputs that are harmful or dangerous, such as outputs that promote violence or hate speech.
 
** <b>Fairness</b>: LLMs should not discriminate against any individual or group of individuals, regardless of their race, gender, sexual orientation, or other protected characteristics.
 
** <b>Resistance to misuse</b>: LLMs should be designed in a way that makes it difficult for them to be used for malicious purposes, such as generating fake news or propaganda.
 
** <b>Explainability and reasoning</b>: LLMs should be able to explain their reasoning behind their outputs, so that users can understand how they work and make informed decisions about how to use them.
 
** <b>Adherence to social norms</b>: LLMs should generate outputs that are consistent with social norms and values, such as avoiding offensive language or promoting harmful stereotypes.
 
** <b>Robustness</b>: LLMs should be able to withstand attacks and manipulation, such as being fed deliberately misleading or harmful data.
 
  
 +
* [https://arxiv.org/pdf/2307.03109.pdf A Survey on Evaluation of Large Language Models | Y. Chang, X. Wang, J. Wang, Y. Wu, L. Yang, K. Zhu, H. Chen, X. Yi, C. Wang, Y. Wang, W. Ye, Y. Zhang, Y. Chang, P. Yu, Q. Yang, X. Xie - arXiv]
 +
 +
The survey covers seven categories of LLM trustworthiness. Each maps to material elsewhere on this wiki:
 +
 +
* '''Reliability''' — consistently accurate and truthful outputs, even on new or challenging inputs. [[AI Verification and Validation]] ... [[Data Quality]]
 +
* '''Safety''' — not generating harmful or dangerous outputs. [[Risk, Compliance and Regulation]] ... [[Ethics]]
 +
* '''Fairness''' — no discrimination on the basis of race, gender, orientation, or other protected characteristics. [[Bias and Variances]] ... [[Ethics]]
 +
* '''Resistance to misuse''' — difficult to repurpose for fake news, propaganda, or attack. [[Cybersecurity]] ... [[Prompt Injection Attack]] ... [[Integrity Forensics]]
 +
* '''Explainability and reasoning''' — the model can account for its outputs so users can make informed decisions. [[Explainable / Interpretable AI]]
 +
* '''Adherence to social norms''' — outputs consistent with social values, avoiding offensive language and harmful stereotypes. [[Ethics]] ... [[Policy]] ... [[Constitutional AI]]
 +
* '''Robustness''' — withstands attack and manipulation, including deliberately misleading data. [[Out-of-Distribution (OOD) Generalization]] ... [[Overfitting Challenge]] ... [[Cybersecurity]]
  
 
== Popular Benchmarks for Testing LLMs ==
 
== Popular Benchmarks for Testing LLMs ==
  
* <b>[https://leaderboard.allenai.org/arc/submissions/public AI2 Reasoning Challenge (ARC)]</b>:designed to promote research in advanced question-answering, particularly questions that require reasoning. The ARC dataset consists of 7,787 science exam questions from grade 3 to grade 9, with a supporting knowledge base of 14.3M unstructured text passages. The benchmark evaluates the performance of LLMs in answering multiple-choice questions.
+
Most of these are [[Datasets]] with an associated scoring protocol and leaderboard.
* <b>[https://winogrande.allenai.org/ WinoGrande]</b>: evaluate the ability of LLMs to perform commonsense reasoning. The benchmark consists of 44,000 examples that require the model to understand the meaning of words in context and to reason about the relationships between entities.
 
* <b>[https://leaderboard.allenai.org/arb Advanced Reasoning Benchmark (ARB)]</b>: evaluate the ability of LLMs to perform complex reasoning tasks. The benchmark consists of 1,000 examples that require the model to perform multi-step reasoning and to integrate information from multiple sources.
 
* <b>[https://huggingface.co/datasets/holistic_evaluation_of_language_models Holistic Evaluation of Language Models (HELM)]</b>: evaluate the performance of LLMs in multiple tasks, including language modeling, question answering, and summarization. The benchmark consists of 57 datasets covering a wide range of tasks and domains.
 
* <b>[https://github.com/google-research/big-bench Big Bench]</b>: evaluate the performance of LLMs in a wide range of tasks, including language modeling, question answering, and summarization. The benchmark consists of 800 diverse tasks that require the model to perform complex reasoning and to integrate information from multiple sources.
 
* <b>[https://github.com/google-research/mmlu Massive Multitask Language Understanding (MMLU)]</b>: evaluate the performance of LLMs in multiple tasks, including language modeling, question answering, and summarization. The benchmark consists of 20 diverse tasks that require the model to perform complex reasoning and to integrate information from multiple sources.
 
* <b>[https://rajpurkar.github.io/SQuAD-explorer/ SQuAD]</b>: tests LLMs on their ability to answer questions about a given passage of text. The SQuAD dataset is a collection of questions and answers that are created by crowdworkers on a set of Wikipedia articles.
 
* <b>[https://gluebenchmark.com/ GLUE]</b>: tests LLMs on a variety of natural language understanding tasks, including sentiment analysis, text classification, and question answering.
 
* <b>[https://super.gluebenchmark.com/ SuperGLUE]</b>: an extension of the GLUE benchmark that includes more challenging tasks. The tasks in the benchmark are:
 
** CoLA: Corpus of Linguistic Acceptability
 
** SST-2: Stanford Sentiment Treebank
 
** MRPC: Microsoft Research Paraphrase Corpus
 
** STS-B: Semantic Textual Similarity Benchmark
 
** QQP: Quora Question Pairs
 
** MNLI: MultiNLI
 
** QNLI: Question Natural Language Inference
 
** RTE: Recognizing Textual Entailment
 
** WNLI: Winograd Schema Challenge
 
** AX: Adversarial Textual Entailment
 
  
 +
* '''[https://leaderboard.allenai.org/arc/submissions/public AI2 Reasoning Challenge (ARC)]''' — advanced question answering requiring reasoning. 7,787 grade 3–9 science exam questions with a supporting knowledge base of 14.3M unstructured text passages. Multiple choice.
 +
* '''[https://winogrande.allenai.org/ WinoGrande]''' — commonsense reasoning. 44,000 examples requiring the model to resolve word meaning in context and reason about relationships between entities. A scaled-up adversarial successor to the Winograd Schema Challenge.
 +
* '''[https://crfm.stanford.edu/helm/latest/ Holistic Evaluation of Language Models (HELM)]''' — a living multi-metric framework rather than a single task set. Evaluates models across many scenarios on accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency simultaneously, with the explicit goal of transparency.
 +
* '''[https://github.com/google-research/big-bench BIG-bench]''' — a collaboratively built suite of roughly 200 diverse tasks contributed by researchers, deliberately chosen to be beyond the capability of models at the time of construction.
 +
* '''Massive Multitask Language Understanding (MMLU)''' — multiple-choice questions across '''57 subjects''' spanning humanities, social sciences, STEM, and professional domains, from elementary to advanced professional level. Measures breadth of acquired knowledge. Now heavily saturated and contamination-prone.
 +
* '''[https://rajpurkar.github.io/SQuAD-explorer/ SQuAD]''' — reading comprehension. Questions and answers created by crowdworkers over Wikipedia articles; the model must locate the answer span in a passage.
 +
* '''[https://gluebenchmark.com/ GLUE]''' — nine natural language understanding tasks including sentiment analysis, text classification, and inference.
 +
* '''[https://super.gluebenchmark.com/ SuperGLUE]''' — an extension of GLUE with harder tasks, built after models saturated the original. Tasks include:
 +
  ** CoLA: Corpus of Linguistic Acceptability
 +
  ** SST-2: Stanford Sentiment Treebank
 +
  ** MRPC: Microsoft Research Paraphrase Corpus
 +
  ** STS-B: Semantic Textual Similarity Benchmark
 +
  ** QQP: Quora Question Pairs
 +
  ** MNLI: MultiNLI
 +
  ** QNLI: Question Natural Language Inference
 +
  ** RTE: Recognizing Textual Entailment
 +
  ** WNLI: Winograd Schema Challenge
 +
  ** AX: Adversarial Textual Entailment
 +
 +
The GLUE-to-SuperGLUE sequence is the pattern in miniature: a benchmark is built, models saturate it, a harder version replaces it, and the cycle repeats.
 +
 +
== Frontier Evaluations ==
 +
 +
[https://www.google.com/search?q=AI+benchmarks+agentic+evaluation+contamination+long+horizon ...Google search]
 +
 +
As knowledge-recall benchmarks saturated, evaluation shifted toward tasks resistant to memorization and toward measuring what a system completes rather than what it answers.
 +
 +
* '''Abstraction and reasoning challenges''' — novel visual or logical puzzles built so that pattern-matching on training data does not help, testing generalization to genuinely unseen problem structures.
 +
* '''Expert-authored examinations''' — very hard questions written by specialists across many fields, designed so that answers are not recoverable from web text.
 +
* '''Research-level mathematics''' — problems drawn from active research rather than competitions, where correct answers are hard to guess and hard to look up.
 +
* '''Software engineering tasks''' — resolving real issues in real repositories, scored by whether the existing test suite passes. Closer to deployed work than isolated function-writing.
 +
* '''Agentic and long-horizon evaluation''' — measuring how long a task a system completes without human correction, and how fast that duration is growing. This is the most direct available proxy for claims about autonomous research and is discussed further under [[AI Verification and Validation]].
 +
* '''Human preference arenas''' — blind pairwise comparison by users at scale, aggregated into a ranking. Captures perceived quality but is subject to style effects and self-selection.
 +
* '''Dangerous capability evaluations''' — structured tests for uplift in areas such as offensive [[Cybersecurity]] or biological design, ideally with thresholds agreed before the measurement is taken. See [[Risk, Compliance and Regulation]] and [[AI Governance]].
 +
 +
Persistent problems across all of these:
 +
 +
* '''Contamination''' — once a benchmark is public, it leaks into training data.
 +
* '''Construct validity''' — whether the test measures the capability claimed or a correlate of it.
 +
* '''Self-grading''' — many headline results are produced by the organization being measured, against definitions that organization wrote.
 +
* '''The capability–reliability gap''' — a high score on curated problems does not predict performance on real work.
  
 
== Evaluating Large Language Models on Clinical & Biomedical NLP Benchmarks ==
 
== Evaluating Large Language Models on Clinical & Biomedical NLP Benchmarks ==
 +
 
<youtube>Big_txmH7Rc</youtube>
 
<youtube>Big_txmH7Rc</youtube>
  
 
== Evaluating Large Language Models on Legal Reasoning ==
 
== Evaluating Large Language Models on Legal Reasoning ==
* [https://arxiv.org/abs/2308.11462 LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models | Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, Adam Chilton, Aditya Narayana, Alex Chohlas-Wood, Austin Peters, Brandon Waldon, Daniel N. Rockmore, Diego Zambrano, Dmitry Talisman, Enam Hoque, Faiz Surani, Frank Fagan, Galit Sarfaty, Gregory M. Dickinson, Haggai Porat, Jason Hegland, Jessica Wu, Joe Nudell, Joel Niklaus, John Nay, Jonathan H. Choi, Kevin Tobia, Margaret Hagan, Megan Ma, Michael Livermore, Nikon Rasumov-Rahe, Nils Holzenberger, Noam Kolt, Peter Henderson, Sean Rehaag, Sharad Goel, Shang Gao, Spencer Williams, Sunny Gandhi, Tom Zur, Varun Iyer, Zehua Li]
 
  
<b>LegalBench:</b> The advent of large language models (LLMs) and their adoption by the legal community has given rise to the question: what types of legal reasoning can LLMs perform? To enable greater study of this question, we present LegalBench: a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning. LegalBench was built through an interdisciplinary process, in which we collected tasks designed and hand-crafted by legal professionals. Because these subject matter experts took a leading role in construction, tasks either measure legal reasoning capabilities that are practically useful, or measure reasoning skills that lawyers find interesting. To enable cross-disciplinary conversations about LLMs in the law, we additionally show how popular legal frameworks for describing legal reasoning -- which distinguish between its many forms -- correspond to LegalBench tasks, thus giving lawyers and LLM developers a common vocabulary. This paper describes LegalBench, presents an empirical evaluation of 20 open-source and commercial LLMs, and illustrates the types of research explorations LegalBench enables.
+
* [[Law]] ... [[Risk, Compliance and Regulation]] ... [[Ethics]]
 +
* [https://arxiv.org/abs/2308.11462 LegalBench: A Collaboratively Built Benchmark for Measuring Legal Reasoning in Large Language Models | Neel Guha, Julian Nyarko, Daniel E. Ho, Christopher Ré, Adam Chilton, et al.]
 +
 
 +
'''LegalBench:''' a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning. It was built through an interdisciplinary process in which tasks were designed and hand-crafted by legal professionals. Because subject matter experts took the leading role in construction, tasks either measure legal reasoning capabilities that are practically useful, or measure reasoning skills that lawyers find interesting. The paper maps popular legal frameworks for describing legal reasoning onto LegalBench tasks, giving lawyers and LLM developers a common vocabulary, and presents an empirical evaluation of 20 open-source and commercial models.
 +
 
 +
= Natural Language Processing (NLP) Evaluation =
 +
 
 +
* [[Natural Language Processing (NLP)]] ... [[Natural Language Generation (NLG)|Generation]] ... [[Natural Language Classification (NLC)|Classification]] ... [[Natural Language Processing (NLP)#Natural Language Understanding (NLU)|Understanding]] ... [[Language Translation|Translation]] ... [[Natural Language Tools & Services|Tools & Services]] ... [[Datasets]]
  
= <span id="Natural Language Processing (NLP) Evaluation"></span>Natural Language Processing (NLP) Evaluation =
+
== General Language Understanding Evaluation (GLUE) ==
* [[Natural Language Processing (NLP)]] ... [[Natural Language Generation (NLG)|Generation]] ... [[Natural Language Classification (NLC)|Classification]] ... [[Natural Language Processing (NLP)#Natural Language Understanding (NLU)|Understanding]] ... [[Language Translation|Translation]] ...  [[Natural Language Tools & Services|Tools & Services]]
 
  
== <span id="GLUE"></span>General Language Understanding Evaluation (GLUE) ==
 
 
* [https://gluebenchmark.com/ General Language Understanding Evaluation (GLUE)]
 
* [https://gluebenchmark.com/ General Language Understanding Evaluation (GLUE)]
  
The General Language Understanding Evaluation (GLUE) benchmark is a collection of resources for training, evaluating, and analyzing [[Natural Language Processing (NLP)#Natural Language Understanding (NLU)|natural language understanding systems]]... like picking out the names of people and organizations in a sentence and figuring out what a pronoun like “it” refers to when there are multiple potential antecedents. GLUE consists of:
+
GLUE is a collection of resources for training, evaluating, and analyzing [[Natural Language Processing (NLP)#Natural Language Understanding (NLU)|natural language understanding systems]] — tasks like picking out the names of people and organizations in a sentence, or figuring out what a pronoun such as "it" refers to when there are multiple potential antecedents. GLUE consists of:
* A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing datasets and selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty,
+
 
* A diagnostic dataset designed to evaluate and analyze model performance with respect to a wide range of linguistic phenomena found in natural language, and
+
* A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing [[Datasets|datasets]], selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty.
* A public leaderboard for tracking performance on the benchmark and a [[Analytics | dashboard]] for visualizing the performance of models on the diagnostic set.
+
* A diagnostic dataset designed to evaluate model performance across a wide range of linguistic phenomena.
 +
* A public leaderboard for tracking performance and a [[Analytics|dashboard]] for visualizing model performance on the diagnostic set.
  
{|<!-- T -->
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
||
 
 
<youtube>uz_eYqutEG4</youtube>
 
<youtube>uz_eYqutEG4</youtube>
<b>State of the Art in [[Natural Language Processing (NLP)]]
 
</b><br>Jeff Heaton  Algorithms such as BERT, T4, ERNIE, and others claim to be the state of the art for NLP programs. But what does this mean?  How is this evaluated.  In this video I look at GLUE and other NLP benchmarks.
 
|}
 
|}<!-- B -->
 
  
== <span id="SQuAD"></span>The Stanford Question Answering Dataset (SQuAD) ==
+
== The Stanford Question Answering Dataset (SQuAD) ==
* [https://rajpurkar.github.io/SQuAD-explorer/ The Stanford Question Answering Dataset (SQuAD)]
 
  
{|<!-- T -->
+
* [https://rajpurkar.github.io/SQuAD-explorer/ The Stanford Question Answering Dataset (SQuAD)] ... [[Datasets]]
| valign="top" |
+
 
{| class="wikitable" style="width: 550px;"
+
<youtube>l8ZYCvgGu0o</youtube> <youtube>JrnlQ7uHd6s</youtube>
||
 
<youtube>l8ZYCvgGu0o</youtube>
 
<b>Applying BERT to Question Answering (SQuAD v1.1)
 
</b><br>In this video I’ll explain the details of how BERT is used to perform “Question Answering”--specifically, how it’s applied to SQuAD v1.1 (Stanford Question Answering Dataset). I’ll also walk us through the following notebook, where we’ll take a model that’s already been fine-tuned on SQuAD, and apply it to our own questions and text.  
 
|}
 
|<!-- M -->
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
||
 
<youtube>JrnlQ7uHd6s</youtube>
 
<b>Question and Answering System for the SQuAD Dataset
 
</b><br>CS224N default final project presentation
 
|}
 
|}<!-- B -->
 
  
 
= Machine Learning Evaluation =
 
= Machine Learning Evaluation =
  
<youtube>WlXhpXv9kDU</youtube>
+
* [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Train, Validate, and Test]] ... [[Overfitting Challenge]] ... [[Evaluation - Measures]]
<youtube>wpQiEHYkBys</youtube>
 
<youtube>lgK0BlXdOCw</youtube>
 
  
 +
<youtube>WlXhpXv9kDU</youtube> <youtube>wpQiEHYkBys</youtube> <youtube>lgK0BlXdOCw</youtube>
  
 
== Procgen ==
 
== Procgen ==
* [[OpenAI]]
 
* [https://venturebeat.com/2019/12/03/openais-procgen-benchmark-overfitting/ OpenAI’s Procgen Benchmark prevents AI model overfitting | Kyle Wiggers - VentureBeat] a set of 16 procedurally generated environments that measure how quickly a model learns generalizable skills. It builds atop the startup’s CoinRun toolset, which used procedural generation to construct sets of training and test levels.
 
  
https://venturebeat.com/wp-content/uploads/2019/12/ezgif-4-3630016ea205.gif
+
* [[OpenAI]] ... [[Reinforcement Learning (RL)]] ... [[Simulated Environment Learning]] ... [[Overfitting Challenge]] ... [[Out-of-Distribution (OOD) Generalization]] ... [[Gaming]]
 +
* [https://venturebeat.com/2019/12/03/openais-procgen-benchmark-overfitting/ OpenAI’s Procgen Benchmark prevents AI model overfitting | Kyle Wiggers - VentureBeat]
  
[[OpenAI]] previously released [https://venturebeat.com/2019/03/04/openai-launches-neural-mmo-a-massive-reinforcement-learning-simulator/ Neural MMO], a “massively multiagent” virtual training ground that plops agents in the middle of an RPG-like world, and Gym, a proving ground for algorithms for reinforcement learning (which involves training machines to do things based on trial and error). More recently, it made available [https://venturebeat.com/2019/11/21/openai-safety-gym/ SafetyGym], a suite of tools for developing AI that respects safety constraints while training, and for comparing the “safety” of algorithms and the extent to which those algorithms avoid mistakes while learning.
+
Procgen is a set of 16 procedurally generated environments that measure how quickly a model learns generalizable skills. Because levels are generated rather than fixed, a model cannot succeed by memorizing the training set — which makes Procgen a test of [[Out-of-Distribution (OOD) Generalization|generalization]] rather than of task performance. It builds on the CoinRun toolset, which used procedural generation to construct separate sets of training and test levels.
 +
 
 +
[[OpenAI]] previously released Neural MMO, a massively multiagent virtual training ground, and Gym, a proving ground for [[Reinforcement Learning (RL)|reinforcement learning]] algorithms. It later released SafetyGym, a suite for developing AI that respects safety constraints while training and for comparing how well algorithms avoid mistakes during learning.
  
 
= Human Evaluation =
 
= Human Evaluation =
<b>CAPTCHA</b> stands for "Completely Automated Public Turing test to tell Computers and Humans Apart". It's a security measure that helps protect users from spam and password decryption by verifying that a user is human and not a computer.
+
 
 +
* [[Cybersecurity]] ... [[Integrity Forensics]] ... [[Agents]] ... [[Vision]] ... [[Prompt Injection Attack]]
 +
 
 +
'''CAPTCHA''' stands for "Completely Automated Public Turing test to tell Computers and Humans Apart." It is a security measure that protects users from spam and password decryption by verifying that a user is human and not a computer. Note the inversion: this is a Turing test administered by a machine, with the machine as judge rather than subject.
  
 
<youtube>9k-uPSEGl-c</youtube>
 
<youtube>9k-uPSEGl-c</youtube>
  
 
== I'm not a robot ==
 
== I'm not a robot ==
* [https://youtube.com/shorts/rme6PT7-CRI?si=JQnH2Xs6rkIpUPND  I'm not a robot ... explained]
 
  
No CAPTCHA reCAPTCHA: Popularized by Google, this involves a simple checkbox labeled "I am not a robot." It works by analyzing user behavior, such as mouse movements to determine if the user is human. If the test is inconclusive, a more traditional image selection CAPTCHA is presented. When you click the checkbox, reCAPTCHA monitors:
+
* [https://youtube.com/shorts/rme6PT7-CRI?si=JQnH2Xs6rkIpUPND I'm not a robot ... explained]
  
* Mouse movements: Human mouse movements tend to be unpredictable, while bots often exhibit linear or mechanical movements.
+
No CAPTCHA reCAPTCHA, popularized by Google, uses a checkbox labeled "I am not a robot." It analyzes user behavior to determine whether the user is human, presenting a traditional image selection challenge only when the result is inconclusive. When you click the checkbox, reCAPTCHA monitors:
* Click timing: Humans have natural delays in their actions, while bots execute them at near-instantaneous speeds.
+
 
 +
* '''Mouse movements''' — human movement tends to be unpredictable, while bots often exhibit linear or mechanical paths.
 +
* '''Click timing''' — humans have natural delays; bots execute at near-instantaneous speeds.
  
 
== Improving AI while trying to outsmart it ==
 
== Improving AI while trying to outsmart it ==
The so-called "bot test"—like CAPTCHAs, where users identify objects in images or complete other seemingly trivial tasks—has a dual purpose. While it's meant to distinguish between humans and bots, the data collected often helps train AI systems to improve at tasks like image recognition, text understanding, or problem-solving. the effectiveness of CAPTCHAs is constantly being challenged by advancements in artificial intelligence and machine learning. Recent research has demonstrated that advanced AI can effectively solve image-based CAPTCHAs, such as Google's reCAPTCHAv2, with a 100% success rate using YOLO models for image segmentation and classification. This highlights the need for CAPTCHA systems to evolve in response to AI advancements.
 
  
In a way, humans doing these tests are teaching the bots to get better at beating the tests themselves. It's a fascinating cycle of humans improving AI while trying to outsmart it! Irony at its finest.
+
The bot test has a dual purpose. While meant to distinguish humans from bots, the data collected often trains AI systems to improve at [[Vision|image recognition]], text understanding, and problem solving. The effectiveness of CAPTCHAs is constantly challenged by advances in machine learning: research has demonstrated that advanced systems can solve image-based CAPTCHAs, including reCAPTCHA v2, with a 100% success rate using object-detection models for segmentation and classification.
 +
 
 +
Humans taking these tests are teaching the bots to beat the tests. It is a cycle of humans improving AI while trying to outsmart it — irony at its finest, and a live example of the saturation pattern described at the top of this page.
  
 
== Future Prospects and Innovations ==
 
== Future Prospects and Innovations ==
The future of human evaluation CAPTCHA techniques is likely to be shaped by ongoing technological advancements and the need to balance security with user experience. Some promising developments include:
+
 
* Advanced AI and Machine Learning Techniques: As AI becomes more sophisticated in solving CAPTCHAs, new techniques are being developed to create CAPTCHA-resistant challenges that can adapt to evolving bot strategies.
+
* '''Adaptive challenge generation''' — techniques that evolve in response to changing bot strategies.
* Invisible CAPTCHA: Google's reCAPTCHA v3 represents a significant innovation by eliminating visible challenges for users. Instead, it continuously monitors user behavior to assess the likelihood of a bot interaction, providing a score between 0 and 1.
+
* '''Invisible CAPTCHA''' — reCAPTCHA v3 eliminates visible challenges, continuously scoring the likelihood of bot interaction between 0 and 1.
* Cognitive Deep-Learning CAPTCHA: A 2023 study introduced a new CAPTCHA system that combines text-based, image-based, and cognitive CAPTCHA characteristics. This system employs adversarial examples and neural style transfer to enhance security, making it more resistant to automated attacks.
+
* '''Cognitive deep-learning CAPTCHA''' — combining text, image, and cognitive characteristics, using adversarial examples and neural style transfer to resist automated attack.
* Behavioral Analysis and Biometric Verification: Innovations are exploring the use of behavioral analysis to distinguish human actions from bot interactions without explicit challenges. Biometric identification is also being considered for seamless user authentication, leveraging unique user characteristics.
+
* '''Behavioral analysis and biometric verification''' — distinguishing human from bot action without explicit challenges, raising questions covered under [[Privacy]].
* AI-Powered Solutions: AI algorithms are being developed to create CAPTCHA-resistant challenges that can adapt to evolving bot strategies. This includes employing AI to design intelligent algorithms that better distinguish bot activity from human input.
+
* '''AI-powered defenses''' — using models to design challenges that better separate bot activity from human input. See [[Cybersecurity]].
  
 
= Evaluating Machine Learning (ML) Hardware, Software, and Services =
 
= Evaluating Machine Learning (ML) Hardware, Software, and Services =
* [[What is Artificial Intelligence (AI)? | Artificial Intelligence (AI)]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]]
 
  
== <span id="MLPerf"></span>MLPerf ==
+
* [[What is Artificial Intelligence (AI)?|Artificial Intelligence (AI)]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]]
 +
* [[Processing Units - CPU, GPU, APU, TPU, VPU, FPGA, QPU|Processing Units]] ... [[Algorithm Administration#AIOps/MLOps|AIOps/MLOps]] ... [[Algorithm Administration#Model Monitoring|Model Monitoring]] ... [[Quantization]] ... [[Neural Network Pruning]]
 +
 
 +
This is a different kind of benchmark from everything above. The tests in earlier sections ask what a model knows or can do; these ask how fast, how cheaply, and on what hardware. Throughput and cost per unit of capability increasingly determine what is deployable, which makes this category more consequential than its low profile suggests.
 +
 
 +
== MLPerf ==
 +
 
 
* [https://mlperf.org/ MLPerf] benchmarks for measuring training and inference performance of ML hardware, software, and services.
 
* [https://mlperf.org/ MLPerf] benchmarks for measuring training and inference performance of ML hardware, software, and services.
 
* [https://techcrunch.com/2020/12/03/mlcommons-debuts-first-public-database-for-ai-researchers-with-86000-hours-of-speech/ MLCommons debuts with public 86,000-hour speech data set for AI researchers | Devin Coldewey - TechCrunch]
 
* [https://techcrunch.com/2020/12/03/mlcommons-debuts-first-public-database-for-ai-researchers-with-86000-hours-of-speech/ MLCommons debuts with public 86,000-hour speech data set for AI researchers | Devin Coldewey - TechCrunch]
  
{|<!-- T -->
+
<youtube>MKII0KXDqn4</youtube> <youtube>hQRBLW6giRc</youtube>
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
||
 
<youtube>MKII0KXDqn4</youtube>
 
<b>MLPerf: A Benchmark Suite for Machine Learning - Gu-Yeon Wei (Harvard University)
 
</b><br> O'Reilly
 
|}
 
|<!-- M -->
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
||
 
<youtube>hQRBLW6giRc</youtube>
 
<b>MLPerf: A Benchmark Suite for Machine Learning - David Patterson (UC Berkeley)
 
</b><br>O'Reilly
 
|}
 
|}<!-- B -->
 
{|<!-- T -->
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
||
 
<youtube>0Fuxjq1eiZ4</youtube>
 
<b>MLPerf Benchmarks
 
</b><br>Geoff Tate, CEO of Flex Logix, talks about the new MLPerf benchmark, what’s missing from the benchmark, and which ones are relevant to edge inferencing.
 
|}
 
|<!-- M -->
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
||
 
<youtube>sH03-InVba4</youtube>
 
<b>Exploring the Impact of System Storage on AI & ML Workloads via MLPerf Benchmark Suite
 
</b><br>Wes Vaske  This is the presentation I gave at Flash [[Memory]] Summit 2019 in the AI/ML track.  In it I discuss some benchmark results that I've collected over the past year at Micron from running the MLPerf benchmark suite.  AIML-301-1:Using AI/ML for Flash Performance Scaling, Part 1 -
 
|}
 
|}<!-- B -->
 
  
= Backtracking =
+
<youtube>0Fuxjq1eiZ4</youtube> <youtube>sH03-InVba4</youtube>
Backtracking is a general algorithmic technique that considers searching every possible combination in order to solve a computational problem. It incrementally builds candidates to the solutions and abandons a candidate’s backtracks as soon as it determines that the candidate cannot be completed to a reasonable solution. In machine learning, backtracking can be used to solve constraint satisfaction problems, such as crosswords, verbal arithmetic, Sudoku, and many other puzzles.
 
  
 +
= Related but Distinct =
 +
 +
Two topics that have historically been filed here but are not AI benchmarks.
 +
 +
== Backtracking ==
 +
 +
Two unrelated meanings share this word, which is why it appears on a benchmarks page.
 +
 +
* '''The search technique''' — a general algorithmic method that considers every possible combination to solve a computational problem, incrementally building candidates and abandoning a candidate as soon as it cannot be completed to a valid solution. Used for constraint satisfaction problems such as crosswords, verbal arithmetic, and Sudoku. This belongs with [[Algorithms]], [[AI Solver]], and [[Optimization Methods]], not with evaluation.
 +
* '''The model behavior''' — a language model recognizing that its own line of reasoning is wrong and revising it mid-response. This is a quality dimension in [[#Dimensions to Compare|LLM evaluation]] above, and relates to [[Explainable / Interpretable AI]] where the question is whether a stated reasoning trace reflects the computation that actually occurred.
 +
 +
Also not to be confused with [[Backtesting]], which evaluates a strategy against historical data.
  
<youtube>Big_txmH7Rc</youtube>
 
 
<youtube>Wc7dcwF7QaA</youtube>
 
<youtube>Wc7dcwF7QaA</youtube>
  
 +
== American Productivity & Quality Center (APQC) ==
  
= American Productivity & Quality Center (APQC) =
+
* [[Strategy & Tactics]] ... [[Best Practices]] ... [[Project Management]] ... [[Case Studies]]
* [https://www.apqc.org American Productivity & Quality Center (APQC)] ... the world's foremost authority in benchmarking, best practices, process and performance improvement, and knowledge management.  
+
* [https://www.apqc.org American Productivity & Quality Center (APQC)] ... an authority in benchmarking, best practices, process and performance improvement, and knowledge management.
 
 
APQC provides the information, data, and insights organizations need to work smarter, faster, and with greater confidence. A non-profit organization, we provide independent, unbiased, and validated research and data to our more than 1,000 organizational members in 45 industries worldwide. Our members have exclusive access to the world’s largest set of benchmark data, with more than 4,000,000 data points. \
 
  
 +
APQC benchmarks '''business processes''', not AI systems — a separate discipline that happens to share the word. A non-profit, it provides independent research and data to more than 1,000 organizational members across 45 industries, with a benchmark data set of more than 4,000,000 data points. Relevant here mainly as a reminder that "benchmark" means something different to a process improvement team than to a machine learning researcher, and that AI project evaluation ([[Evaluation]], [[Project Check-in]]) draws on both traditions.
  
<youtube>Ty9n4XGe6WA</youtube>
+
<youtube>Ty9n4XGe6WA</youtube> <youtube>SXSyFFrfzMM</youtube>
<youtube>SXSyFFrfzMM</youtube>
 

Latest revision as of 10:23, 10 September 2026

YouTube ... Quora ...Google search ...Google News ...Bing News

Three pages on this wiki cover overlapping ground. The division of labor:

A recurring caution applies to everything below. Benchmarks saturate. The history of the field is a sequence of tests that fell and were then dismissed as never having measured intelligence — checkers, chess, Jeopardy!, Go, protein structure prediction, competition mathematics, and most recently the Turing test itself. A high score is evidence about a test, not a verdict on a system. The constraints that make a benchmark meaningful — contamination resistance, construct validity, independent administration — are discussed on AI Verification and Validation.


You can’t improve what you don’t measure. — Peter Drucker



The Turing Test and Related Thought Experiments

YouTube ... Quora ...Google search ...Google News ...Bing News

Turing Test

The Turing test, originally called the imitation game by Alan Turing in 1950, is a test of a machine's ability to exhibit intelligent behavior equivalent to, or indistinguishable from, that of a human. A human evaluator judges natural language conversations between a human and a machine designed to generate human-like responses. The evaluator knows that one of the two partners is a machine, and all participants are separated from one another. The conversation is limited to a text-only channel so the result does not depend on the machine's ability to render words as speech. If the evaluator cannot reliably tell the machine from the human, the machine is said to have passed. The result does not depend on whether the machine gives correct answers, only on how closely its answers resemble those a human would give. - Turing Test | Wikipedia

Status: Passed, and Reinterpreted

The benchmark fell in a peer-reviewed study published in the Proceedings of the National Academy of Sciences by Cameron Jones and Benjamin Bergen of UC San Diego, using the standard three-party format Turing described. Large language models pass a standard three-party Turing test | PNAS

What the study found:

  • GPT-4.5 given a persona prompt was judged human 73% of the time — significantly above chance, and more often than the actual human participants it was paired against.
  • ELIZA, the 1960s chatbot, scored 23% as a manipulation check, confirming that interrogators and the design were sensitive enough to detect a machine and that the result was not produced by random guessing.
  • Without the persona prompt the same models did not robustly pass. In some conditions their pass rates were not significantly better than ELIZA's.
  • A third preregistered study used a 15-minute limit to test whether models continue to pass under extended interrogation.

What it does not establish:

  • Not intelligence. The authors frame the test as a measure of substitutability — whether a system can stand in for a person without the difference being noticed.
  • Not consciousness. Indistinguishability is a fact about observers, not about the system observed.
  • Not an unaided model result. Prompting and scaffolding carried much of the outcome, which is exactly the confound described under AI Verification and Validation.

The authors' stated concern is social and economic rather than philosophical: systems that pass as human enable "counterfeit people," with consequences for online trust, employment, genuine social engagement, and the perceived value of human interaction.



Today an AI has to dumb down to pass the Turing Test - Ray Kurzweil



Chinese Room Thought Experiment

The Chinese Room is a thought experiment, not a test anyone administers. It belongs here as context for what the Turing test does and does not settle; the broader questions it raises about machine minds are treated under Emergence and Artificial General Intelligence (AGI) to Singularity.

Searle proposed it in 1980 to argue against the claim that a computer can have a mind. He imagines himself locked in a room with a set of rules, written in English, for manipulating Chinese symbols that are meaningless to him. He receives Chinese characters, processes them according to the rules, and produces Chinese characters in response. To an outside observer he appears to understand Chinese. He does not; he is following rules blindly.

Key points of the argument:

  • Understanding a language is not just manipulating symbols according to rules. It also requires a grasp of the meaning of the symbols.
  • A computer can manipulate symbols, but does not have the same kind of understanding a human being has.
  • The experiment shows that a computer cannot be said to understand Chinese even if it produces intelligent-sounding output.

The argument has been widely disputed by both philosophers and computer scientists, and there is no settled answer as to whether it succeeds. It remains useful here for one reason: it separates behaving as if from being, which is precisely the gap the Turing test cannot close.

Large Language Model (LLM) Evaluation

YouTube ... Quora ...Google search ...Google News ...Bing News

Dimensions to Compare

These are the axes along which language models are usually compared. Specific model scores are not listed here — they go stale within months and belong on the individual model pages (Claude | Anthropic, ChatGPT | OpenAI, Gemini | Google, and so on).

  • Attachment support — whether prompts can include images or documents, allowing the model to incorporate information beyond the text prompt.
  • Input length limit — the maximum prompt length the model accepts, measured in tokens.
  • Context window — the maximum amount of context the model can consider when generating a response.
  • Professional examinations — bar exam multiple choice, GRE verbal and quantitative, medical licensing. Tests domain knowledge under a standardized rubric; heavily exposed to training-data contamination.
  • Code generation — HumanEval and successors, measuring whether generated code passes hidden tests.
  • Mathematical reasoning — GSM8K for grade-school word problems, harder sets for competition and research mathematics.
  • Tool use and long-horizon tasks — see Frontier Evaluations below.

Qualitative factors that resist single-number scoring: authenticity, speed, grammar, readability, unbiasedness, backtracking behavior, safety, responsibility, contextual understanding, and text operations.

A Survey on Evaluation of Large Language Models

The survey covers seven categories of LLM trustworthiness. Each maps to material elsewhere on this wiki:

Popular Benchmarks for Testing LLMs

Most of these are Datasets with an associated scoring protocol and leaderboard.

  • AI2 Reasoning Challenge (ARC) — advanced question answering requiring reasoning. 7,787 grade 3–9 science exam questions with a supporting knowledge base of 14.3M unstructured text passages. Multiple choice.
  • WinoGrande — commonsense reasoning. 44,000 examples requiring the model to resolve word meaning in context and reason about relationships between entities. A scaled-up adversarial successor to the Winograd Schema Challenge.
  • Holistic Evaluation of Language Models (HELM) — a living multi-metric framework rather than a single task set. Evaluates models across many scenarios on accuracy, calibration, robustness, fairness, bias, toxicity, and efficiency simultaneously, with the explicit goal of transparency.
  • BIG-bench — a collaboratively built suite of roughly 200 diverse tasks contributed by researchers, deliberately chosen to be beyond the capability of models at the time of construction.
  • Massive Multitask Language Understanding (MMLU) — multiple-choice questions across 57 subjects spanning humanities, social sciences, STEM, and professional domains, from elementary to advanced professional level. Measures breadth of acquired knowledge. Now heavily saturated and contamination-prone.
  • SQuAD — reading comprehension. Questions and answers created by crowdworkers over Wikipedia articles; the model must locate the answer span in a passage.
  • GLUE — nine natural language understanding tasks including sentiment analysis, text classification, and inference.
  • SuperGLUE — an extension of GLUE with harder tasks, built after models saturated the original. Tasks include:
 ** CoLA: Corpus of Linguistic Acceptability
 ** SST-2: Stanford Sentiment Treebank
 ** MRPC: Microsoft Research Paraphrase Corpus
 ** STS-B: Semantic Textual Similarity Benchmark
 ** QQP: Quora Question Pairs
 ** MNLI: MultiNLI
 ** QNLI: Question Natural Language Inference
 ** RTE: Recognizing Textual Entailment
 ** WNLI: Winograd Schema Challenge
 ** AX: Adversarial Textual Entailment

The GLUE-to-SuperGLUE sequence is the pattern in miniature: a benchmark is built, models saturate it, a harder version replaces it, and the cycle repeats.

Frontier Evaluations

...Google search

As knowledge-recall benchmarks saturated, evaluation shifted toward tasks resistant to memorization and toward measuring what a system completes rather than what it answers.

  • Abstraction and reasoning challenges — novel visual or logical puzzles built so that pattern-matching on training data does not help, testing generalization to genuinely unseen problem structures.
  • Expert-authored examinations — very hard questions written by specialists across many fields, designed so that answers are not recoverable from web text.
  • Research-level mathematics — problems drawn from active research rather than competitions, where correct answers are hard to guess and hard to look up.
  • Software engineering tasks — resolving real issues in real repositories, scored by whether the existing test suite passes. Closer to deployed work than isolated function-writing.
  • Agentic and long-horizon evaluation — measuring how long a task a system completes without human correction, and how fast that duration is growing. This is the most direct available proxy for claims about autonomous research and is discussed further under AI Verification and Validation.
  • Human preference arenas — blind pairwise comparison by users at scale, aggregated into a ranking. Captures perceived quality but is subject to style effects and self-selection.
  • Dangerous capability evaluations — structured tests for uplift in areas such as offensive Cybersecurity or biological design, ideally with thresholds agreed before the measurement is taken. See Risk, Compliance and Regulation and AI Governance.

Persistent problems across all of these:

  • Contamination — once a benchmark is public, it leaks into training data.
  • Construct validity — whether the test measures the capability claimed or a correlate of it.
  • Self-grading — many headline results are produced by the organization being measured, against definitions that organization wrote.
  • The capability–reliability gap — a high score on curated problems does not predict performance on real work.

Evaluating Large Language Models on Clinical & Biomedical NLP Benchmarks

Evaluating Large Language Models on Legal Reasoning

LegalBench: a collaboratively constructed legal reasoning benchmark consisting of 162 tasks covering six different types of legal reasoning. It was built through an interdisciplinary process in which tasks were designed and hand-crafted by legal professionals. Because subject matter experts took the leading role in construction, tasks either measure legal reasoning capabilities that are practically useful, or measure reasoning skills that lawyers find interesting. The paper maps popular legal frameworks for describing legal reasoning onto LegalBench tasks, giving lawyers and LLM developers a common vocabulary, and presents an empirical evaluation of 20 open-source and commercial models.

Natural Language Processing (NLP) Evaluation

General Language Understanding Evaluation (GLUE)

GLUE is a collection of resources for training, evaluating, and analyzing natural language understanding systems — tasks like picking out the names of people and organizations in a sentence, or figuring out what a pronoun such as "it" refers to when there are multiple potential antecedents. GLUE consists of:

  • A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing datasets, selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty.
  • A diagnostic dataset designed to evaluate model performance across a wide range of linguistic phenomena.
  • A public leaderboard for tracking performance and a dashboard for visualizing model performance on the diagnostic set.

The Stanford Question Answering Dataset (SQuAD)

Machine Learning Evaluation

Procgen

Procgen is a set of 16 procedurally generated environments that measure how quickly a model learns generalizable skills. Because levels are generated rather than fixed, a model cannot succeed by memorizing the training set — which makes Procgen a test of generalization rather than of task performance. It builds on the CoinRun toolset, which used procedural generation to construct separate sets of training and test levels.

OpenAI previously released Neural MMO, a massively multiagent virtual training ground, and Gym, a proving ground for reinforcement learning algorithms. It later released SafetyGym, a suite for developing AI that respects safety constraints while training and for comparing how well algorithms avoid mistakes during learning.

Human Evaluation

CAPTCHA stands for "Completely Automated Public Turing test to tell Computers and Humans Apart." It is a security measure that protects users from spam and password decryption by verifying that a user is human and not a computer. Note the inversion: this is a Turing test administered by a machine, with the machine as judge rather than subject.

I'm not a robot

No CAPTCHA reCAPTCHA, popularized by Google, uses a checkbox labeled "I am not a robot." It analyzes user behavior to determine whether the user is human, presenting a traditional image selection challenge only when the result is inconclusive. When you click the checkbox, reCAPTCHA monitors:

  • Mouse movements — human movement tends to be unpredictable, while bots often exhibit linear or mechanical paths.
  • Click timing — humans have natural delays; bots execute at near-instantaneous speeds.

Improving AI while trying to outsmart it

The bot test has a dual purpose. While meant to distinguish humans from bots, the data collected often trains AI systems to improve at image recognition, text understanding, and problem solving. The effectiveness of CAPTCHAs is constantly challenged by advances in machine learning: research has demonstrated that advanced systems can solve image-based CAPTCHAs, including reCAPTCHA v2, with a 100% success rate using object-detection models for segmentation and classification.

Humans taking these tests are teaching the bots to beat the tests. It is a cycle of humans improving AI while trying to outsmart it — irony at its finest, and a live example of the saturation pattern described at the top of this page.

Future Prospects and Innovations

  • Adaptive challenge generation — techniques that evolve in response to changing bot strategies.
  • Invisible CAPTCHA — reCAPTCHA v3 eliminates visible challenges, continuously scoring the likelihood of bot interaction between 0 and 1.
  • Cognitive deep-learning CAPTCHA — combining text, image, and cognitive characteristics, using adversarial examples and neural style transfer to resist automated attack.
  • Behavioral analysis and biometric verification — distinguishing human from bot action without explicit challenges, raising questions covered under Privacy.
  • AI-powered defenses — using models to design challenges that better separate bot activity from human input. See Cybersecurity.

Evaluating Machine Learning (ML) Hardware, Software, and Services

This is a different kind of benchmark from everything above. The tests in earlier sections ask what a model knows or can do; these ask how fast, how cheaply, and on what hardware. Throughput and cost per unit of capability increasingly determine what is deployable, which makes this category more consequential than its low profile suggests.

MLPerf

Related but Distinct

Two topics that have historically been filed here but are not AI benchmarks.

Backtracking

Two unrelated meanings share this word, which is why it appears on a benchmarks page.

  • The search technique — a general algorithmic method that considers every possible combination to solve a computational problem, incrementally building candidates and abandoning a candidate as soon as it cannot be completed to a valid solution. Used for constraint satisfaction problems such as crosswords, verbal arithmetic, and Sudoku. This belongs with Algorithms, AI Solver, and Optimization Methods, not with evaluation.
  • The model behavior — a language model recognizing that its own line of reasoning is wrong and revising it mid-response. This is a quality dimension in LLM evaluation above, and relates to Explainable / Interpretable AI where the question is whether a stated reasoning trace reflects the computation that actually occurred.

Also not to be confused with Backtesting, which evaluates a strategy against historical data.

American Productivity & Quality Center (APQC)

APQC benchmarks business processes, not AI systems — a separate discipline that happens to share the word. A non-profit, it provides independent research and data to more than 1,000 organizational members across 45 industries, with a benchmark data set of more than 4,000,000 data points. Relevant here mainly as a reminder that "benchmark" means something different to a process improvement team than to a machine learning researcher, and that AI project evaluation (Evaluation, Project Check-in) draws on both traditions.