Difference between revisions of "Benchmarks"
| Line 5: | Line 5: | ||
|description=Helpful resources for your journey with artificial intelligence; videos, articles, techniques, courses, profiles, and tools | |description=Helpful resources for your journey with artificial intelligence; videos, articles, techniques, courses, profiles, and tools | ||
}} | }} | ||
| − | [http://www.youtube.com/results?search_query=Benchmark+machine+learning+ | + | [http://www.youtube.com/results?search_query=Benchmark+machine+learning+model YouTube search...] |
| − | [http://www.google.com/search?q=Benchmark+machine+learning+ | + | [http://www.google.com/search?q=Benchmark+machine+learning+model ...Google search] |
* [[Evaluation Measures - Classification Performance]] - Accuracy, Precision & Recall (Sensitivity), and Specificity | * [[Evaluation Measures - Classification Performance]] - Accuracy, Precision & Recall (Sensitivity), and Specificity | ||
| Line 14: | Line 14: | ||
<youtube>YygGzfkhtJc</youtube> | <youtube>YygGzfkhtJc</youtube> | ||
| − | <youtube> | + | <youtube>hQRBLW6giRc</youtube> |
| − | <youtube> | + | <youtube>WlXhpXv9kDU</youtube> |
<youtube>7fcWfUavO7E</youtube> | <youtube>7fcWfUavO7E</youtube> | ||
Revision as of 10:31, 24 December 2019
YouTube search... ...Google search
- Evaluation Measures - Classification Performance - Accuracy, Precision & Recall (Sensitivity), and Specificity
- Datasets
- DAWNBench | Stanford - an End-to-End Deep Learning Benchmark and Competition
GLUE
The General Language Understanding Evaluation (GLUE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems. GLUE consists of: A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing datasets and selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty, A diagnostic dataset designed to evaluate and analyze model performance with respect to a wide range of linguistic phenomena found in natural language, and A public leaderboard for tracking performance on the benchmark and a dashboard for visualizing the performance of models on the diagnostic set.