Difference between revisions of "Benchmarks"

From
Jump to: navigation, search
(Created page with "{{#seo: |title=PRIMO.ai |titlemode=append |keywords=artificial, intelligence, machine, learning, models, algorithms, data, singularity, moonshot, Tensorflow, Google, Nvidia, M...")
 
Line 8: Line 8:
 
[http://www.google.com/search?q=Benchmark+machine+learning+ML ...Google search]
 
[http://www.google.com/search?q=Benchmark+machine+learning+ML ...Google search]
  
 +
* [[Evaluation Measures - Classification Performance]] - Accuracy, Precision & Recall (Sensitivity), and Specificity
 
* [[Datasets]]
 
* [[Datasets]]
* [http://dawn.cs.stanford.edu//benchmark/index.html DAWNBench] - an End-to-End Deep Learning Benchmark and Competition
+
* [http://dawn.cs.stanford.edu//benchmark/index.html DAWNBench | Stanford] - an End-to-End Deep Learning Benchmark and Competition
  
<youtube>uz_eYqutEG4</youtube>
+
 
 +
<youtube>YygGzfkhtJc</youtube>
 
<youtube>cw2LvVkmtkQ</youtube>
 
<youtube>cw2LvVkmtkQ</youtube>
 
<youtube>TK-2189UcKk</youtube>
 
<youtube>TK-2189UcKk</youtube>
Line 17: Line 19:
  
 
== GLUE ==
 
== GLUE ==
 +
* [http://gluebenchmark.com/ GLUE]
 +
 +
The General Language Understanding Evaluation (GLUE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems. GLUE consists of:
 +
A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing datasets and selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty,
 +
A diagnostic dataset designed to evaluate and analyze model performance with respect to a wide range of linguistic phenomena found in natural language, and
 +
A public leaderboard for tracking performance on the benchmark and a dashboard for visualizing the performance of models on the diagnostic set.
 +
 
<youtube>uz_eYqutEG4</youtube>
 
<youtube>uz_eYqutEG4</youtube>

Revision as of 10:27, 24 December 2019

YouTube search... ...Google search


GLUE

The General Language Understanding Evaluation (GLUE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems. GLUE consists of: A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing datasets and selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty, A diagnostic dataset designed to evaluate and analyze model performance with respect to a wide range of linguistic phenomena found in natural language, and A public leaderboard for tracking performance on the benchmark and a dashboard for visualizing the performance of models on the diagnostic set.