Difference between revisions of "Benchmarks"
(Created page with "{{#seo: |title=PRIMO.ai |titlemode=append |keywords=artificial, intelligence, machine, learning, models, algorithms, data, singularity, moonshot, Tensorflow, Google, Nvidia, M...") |
|||
| Line 8: | Line 8: | ||
[http://www.google.com/search?q=Benchmark+machine+learning+ML ...Google search] | [http://www.google.com/search?q=Benchmark+machine+learning+ML ...Google search] | ||
| + | * [[Evaluation Measures - Classification Performance]] - Accuracy, Precision & Recall (Sensitivity), and Specificity | ||
* [[Datasets]] | * [[Datasets]] | ||
| − | * [http://dawn.cs.stanford.edu//benchmark/index.html DAWNBench] - an End-to-End Deep Learning Benchmark and Competition | + | * [http://dawn.cs.stanford.edu//benchmark/index.html DAWNBench | Stanford] - an End-to-End Deep Learning Benchmark and Competition |
| − | <youtube> | + | |
| + | <youtube>YygGzfkhtJc</youtube> | ||
<youtube>cw2LvVkmtkQ</youtube> | <youtube>cw2LvVkmtkQ</youtube> | ||
<youtube>TK-2189UcKk</youtube> | <youtube>TK-2189UcKk</youtube> | ||
| Line 17: | Line 19: | ||
== GLUE == | == GLUE == | ||
| + | * [http://gluebenchmark.com/ GLUE] | ||
| + | |||
| + | The General Language Understanding Evaluation (GLUE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems. GLUE consists of: | ||
| + | A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing datasets and selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty, | ||
| + | A diagnostic dataset designed to evaluate and analyze model performance with respect to a wide range of linguistic phenomena found in natural language, and | ||
| + | A public leaderboard for tracking performance on the benchmark and a dashboard for visualizing the performance of models on the diagnostic set. | ||
| + | |||
<youtube>uz_eYqutEG4</youtube> | <youtube>uz_eYqutEG4</youtube> | ||
Revision as of 10:27, 24 December 2019
YouTube search... ...Google search
- Evaluation Measures - Classification Performance - Accuracy, Precision & Recall (Sensitivity), and Specificity
- Datasets
- DAWNBench | Stanford - an End-to-End Deep Learning Benchmark and Competition
GLUE
The General Language Understanding Evaluation (GLUE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems. GLUE consists of: A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing datasets and selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty, A diagnostic dataset designed to evaluate and analyze model performance with respect to a wide range of linguistic phenomena found in natural language, and A public leaderboard for tracking performance on the benchmark and a dashboard for visualizing the performance of models on the diagnostic set.