Difference between revisions of "Benchmarks"

From
Jump to: navigation, search
Line 8: Line 8:
 
[http://www.google.com/search?q=Benchmark+machine+learning+model ...Google search]
 
[http://www.google.com/search?q=Benchmark+machine+learning+model ...Google search]
  
* [[Evaluation Measures - Classification Performance]] - Accuracy, Precision & Recall (Sensitivity), and Specificity
+
* [[Evaluation Measures - Classification Performance]] - [[Evaluation Measures - Classification Performance#Accuracy|Accuracy]], [[Evaluation Measures - Classification Performance#Precision & Recall (Sensitivity)|Precision & Recall (Sensitivity)]], and [[Evaluation Measures - Classification Performance#Specificity|Specificity]]
 
* [[Datasets]]
 
* [[Datasets]]
 
* [http://dawn.cs.stanford.edu//benchmark/index.html DAWNBench | Stanford] - an End-to-End Deep Learning Benchmark and Competition
 
* [http://dawn.cs.stanford.edu//benchmark/index.html DAWNBench | Stanford] - an End-to-End Deep Learning Benchmark and Competition
  
 +
https://www.aitrends.com/wp-content/uploads/2018/05/6-1MachineLearning-6.jpg
  
 
<youtube>YygGzfkhtJc</youtube>
 
<youtube>YygGzfkhtJc</youtube>
 
<youtube>hQRBLW6giRc</youtube>
 
<youtube>hQRBLW6giRc</youtube>
 
<youtube>WlXhpXv9kDU</youtube>
 
<youtube>WlXhpXv9kDU</youtube>
<youtube>7fcWfUavO7E</youtube>
+
<youtube>wpQiEHYkBys</youtube>
 +
 
  
 
== GLUE ==
 
== GLUE ==

Revision as of 10:37, 24 December 2019

YouTube search... ...Google search

6-1MachineLearning-6.jpg


GLUE

The General Language Understanding Evaluation (GLUE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems. GLUE consists of: A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing datasets and selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty, A diagnostic dataset designed to evaluate and analyze model performance with respect to a wide range of linguistic phenomena found in natural language, and A public leaderboard for tracking performance on the benchmark and a dashboard for visualizing the performance of models on the diagnostic set.