Difference between revisions of "Benchmarks"

From
Jump to: navigation, search
(Procgen)
(Procgen)
Line 55: Line 55:
 
* [http://venturebeat.com/2019/12/03/openais-procgen-benchmark-overfitting/ OpenAI’s Procgen Benchmark prevents AI model overfitting | Kyle Wiggers - VentureBeat] a set of 16 procedurally generated environments that measure how quickly a model learns generalizable skills. It builds atop the startup’s CoinRun toolset, which used procedural generation to construct sets of training and test levels.
 
* [http://venturebeat.com/2019/12/03/openais-procgen-benchmark-overfitting/ OpenAI’s Procgen Benchmark prevents AI model overfitting | Kyle Wiggers - VentureBeat] a set of 16 procedurally generated environments that measure how quickly a model learns generalizable skills. It builds atop the startup’s CoinRun toolset, which used procedural generation to construct sets of training and test levels.
  
<youtube>9FAXAgRrOSE</youtube>
+
http://venturebeat.com/wp-content/uploads/2019/12/ezgif-4-3630016ea205.gif?w=360&resize=360%2C360&strip=all
<youtube>G5rq5ei2Lec</youtube>
 

Revision as of 23:20, 26 December 2019

YouTube search... ...Google search



GLUE

The General Language Understanding Evaluation (GLUE) benchmark is a collection of resources for training, evaluating, and analyzing natural language understanding systems... like picking out the names of people and organizations in a sentence and figuring out what a pronoun like “it” refers to when there are multiple potential antecedents. GLUE consists of: A benchmark of nine sentence- or sentence-pair language understanding tasks built on established existing datasets and selected to cover a diverse range of dataset sizes, text genres, and degrees of difficulty, A diagnostic dataset designed to evaluate and analyze model performance with respect to a wide range of linguistic phenomena found in natural language, and A public leaderboard for tracking performance on the benchmark and a dashboard for visualizing the performance of models on the diagnostic set.

The Stanford Question Answering Dataset (SQuAD)

ReAding Comprehension (RACE)


MLPerf

  • MLPerf benchmarks for measuring training and inference performance of ML hardware, software, and services.

Procgen

http://venturebeat.com/wp-content/uploads/2019/12/ezgif-4-3630016ea205.gif?w=360&resize=360%2C360&strip=all