Difference between revisions of "Moonshots"

From
Jump to: navigation, search
m
Line 21: Line 21:
 
* [[Humor]]
 
* [[Humor]]
 
** [https://www.vice.com/en/article/z43nke/joke-telling-robots-are-the-final-frontier-of-artificial-intelligence Joke-Telling Robots Are the Final Frontier of Artificial Intelligence | Becky Ferreira - Vice] ... [[Humor]] requires self-awareness, spontaneity, linguistic sophistication, and empathy. Not easy for a [[Robotics | robot]].
 
** [https://www.vice.com/en/article/z43nke/joke-telling-robots-are-the-final-frontier-of-artificial-intelligence Joke-Telling Robots Are the Final Frontier of Artificial Intelligence | Becky Ferreira - Vice] ... [[Humor]] requires self-awareness, spontaneity, linguistic sophistication, and empathy. Not easy for a [[Robotics | robot]].
*[https://medium.com/applied-data-science/how-to-build-your-own-alphazero-ai-using-python-and-keras-7f664945c188 How to build your own AlphaZero AI using Python and Keras]
 
 
* [[What is Artificial Intelligence (AI)? | Artificial Intelligence (AI)]] ... [[Generative AI]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]]
 
* [[What is Artificial Intelligence (AI)? | Artificial Intelligence (AI)]] ... [[Generative AI]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]]
 
* [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[Bing/Copilot]] | [[Microsoft]] ... [[Gemini]] | [[Google]] ... [[Claude]] | [[Anthropic]] ... [[Perplexity]] ... [[You]] ... [[phind]] ... [[Ernie]] | [[Baidu]]
 
* [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[Bing/Copilot]] | [[Microsoft]] ... [[Gemini]] | [[Google]] ... [[Claude]] | [[Anthropic]] ... [[Perplexity]] ... [[You]] ... [[phind]] ... [[Ernie]] | [[Baidu]]
 
* [[Data Science]] ... [[Data Governance|Governance]] ... [[Data Preprocessing|Preprocessing]] ... [[Feature Exploration/Learning|Exploration]] ... [[Data Interoperability|Interoperability]] ... [[Algorithm Administration#Master Data Management (MDM)|Master Data Management (MDM)]] ... [[Bias and Variances]] ... [[Benchmarks]] ... [[Datasets]]  
 
* [[Data Science]] ... [[Data Governance|Governance]] ... [[Data Preprocessing|Preprocessing]] ... [[Feature Exploration/Learning|Exploration]] ... [[Data Interoperability|Interoperability]] ... [[Algorithm Administration#Master Data Management (MDM)|Master Data Management (MDM)]] ... [[Bias and Variances]] ... [[Benchmarks]] ... [[Datasets]]  
 +
*[https://medium.com/applied-data-science/how-to-build-your-own-alphazero-ai-using-python-and-keras-7f664945c188 How to build your own AlphaZero AI using Python and Keras]
 
* [https://www.nature.com/articles/s41586-024-07936-z The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Sakana AI Team - Nature] ...Breakthrough in AI-generated research papers achieving human-level acceptance at workshops.
 
* [https://www.nature.com/articles/s41586-024-07936-z The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Sakana AI Team - Nature] ...Breakthrough in AI-generated research papers achieving human-level acceptance at workshops.
 
* [https://www.congress.gov/bill/119th-congress/senate-bill/AI-Grand-Challenges-Act AI Grand Challenges Act | U.S. Senate - 2026] ...Proposed bipartisan legislation to fund prize competitions for AI breakthroughs in health and security.
 
* [https://www.congress.gov/bill/119th-congress/senate-bill/AI-Grand-Challenges-Act AI Grand Challenges Act | U.S. Senate - 2026] ...Proposed bipartisan legislation to fund prize competitions for AI breakthroughs in health and security.

Revision as of 07:12, 10 September 2026

YouTube search... ... Quora search ...Google search ...Google News ...Bing News


The “Sputnik” moment for China came a year ago when a Google computer program, AlphaGo, beat the world’s top master of the ancient board game of Go. Now, China is racing to become the world leader in artificial-intelligence. The term "moonshot" is derived from the Apollo program, which was a series of space missions undertaken by the United States in the 1960s and early 1970s with the goal of landing humans on the Moon. The Apollo program was considered a moonshot because it represented a major technological and engineering challenge that required significant innovation and investment. In context, what do you think would be a "Moonshot" response?



The 'moonshot' milestones along the road to Artificial General Intelligence (AGI)



In the context of AI, a "moonshot" refers to a project or goal that aims to achieve a major breakthrough in artificial intelligence that has the potential to transform society or address significant global challenges. Moonshots are high-risk, high-impact goals where advanced AI is used to tackle problems that previously seemed decades away—such as autonomous scientific discovery, general-purpose robotics, radical health extension, new energy technologies, and eventually AI systems that improve AI itself.


Able to Predict the Future

Youtube search... ...Google search

Can Conjure & Ask Questions

Youtube search... ...Google search

Recursive Self Improvement

The topic of recursive self-improvement is a significant threshold in AI development. Currently, the process is largely driven by human software engineers who manually generate training data, run ablations to test data quality, and evaluate models against benchmarks.

However, many labs are now working to close this loop to automate the process:

  • Automated Judges: Different models will act as judges to evaluate the quality of outputs.
  • Generative Feedback: Models will generate new, high-quality training data autonomously.
  • Adversarial Reasoning: Advanced models will reason over which data to include, effectively filtering for quality.

By feeding this output back into the post-training process, development speed will likely increase significantly. While some speculate this could lead to an intelligence explosion, Mustafa Suleyman notes that achieving this requires substantial compute and, without proper human oversight or control, it introduces significant risks.


The frontier of AI has shifted from static chat interfaces to autonomous "Agentic Workflows." These systems are designed for recursive self-improvement and multi-agent orchestration, moving beyond simple prompt-response cycles.

The Automated Researcher

By September 2026, OpenAI hit a major milestone with its "automated research intern." Think of it as a tireless assistant that handles well-defined research tasks while a human steers the ship. For every eight hours a person puts in, the system completes about three days' worth of work. The ultimate goal is to have a fully autonomous AI researcher running by March 2028.

AI as Scientist

We have moved past seeing AI as just a tool; it is now a true collaborator. Models like Sakana AI's "AI Scientist" are actually producing fresh, peer-review-quality research. They are running experiments in areas like virtual cell modeling and designing new drugs from scratch. It is like having a digital post-doc working around the clock in the lab.

Unsolved Mathematics & Formal Verification

After hitting gold-medal performance in the International Mathematical Olympiad, AI is setting its sights on creating original proofs for open mathematical problems. To make sure the math is rock-solid, these systems use formal verification tools like Lean. Imagine a spell-checker, but instead of catching typos, it verifies complex logic step-by-step to guarantee accuracy.

Reliability & Alignment

Continual learning and long-term reliability are the unglamorous but necessary hurdles we still need to clear. Right now, current models tend to lose focus or degrade when they run on their own for too long. If we want AI agents to operate stably over weeks or months, fixing this drift is essential. Think of it like maintaining focus during a marathon rather than just running a quick sprint.

Embodied General Intelligence

Many people call this the "physical Turing test." Embodied General Intelligence is all about getting robots to navigate messy, unpredictable real-world spaces. Instead of repeating the same motion on an assembly line, these robots need to figure out how to fold laundry or do the dishes in a kitchen they have never seen before.

Autonomous Vehicles

Youtube search... ...Google search


Need to 'Learn' the Wide World Web

Youtube search... ...Google search

Discussions on the Future of AI

Winograd Schema Challenge (WSC): From AI Moonshot to Historical Milestone

YouTube search... ...Google search

The Winograd Schema Challenge (WSC) was proposed by Hector Levesque in 2011 as a more focused alternative to the Turing Test. It tested whether an AI system could resolve ambiguous pronouns in sentences where the correct answer appears to require commonsense reasoning and knowledge about how the world works.

A classic example is:

The city councilmen refused the demonstrators a permit because they feared violence.
The city councilmen refused the demonstrators a permit because they advocated violence.

In the first sentence, they normally refers to the city councilmen; in the second, it refers to the demonstrators. Humans generally make this distinction effortlessly by using contextual and commonsense knowledge rather than grammatical rules alone.

When the challenge was introduced, solving problems like these was considered a significant goal for AI. The original schemas were deliberately constructed to resist simple statistical shortcuts, making the WSC an influential test of whether machines could move beyond pattern matching toward genuine language understanding and reasoning. "The Winograd Schema Challenge" | Ernest Davis, Leora Morgenstern, and Charles Ortiz

The status of the challenge has since changed. By 2019, multiple systems based on large pretrained transformer language models and fine-tuned on related tasks were achieving greater than 90 percent accuracy. A later review by Vid Kocijan, Ernest Davis, Thomas Lukasiewicz, Gary Marcus, and Leora Morgenstern concluded that the original WSC had largely been overcome as a benchmark. "The Defeat of the Winograd Schema Challenge" | Artificial Intelligence, 2023

This did not mean that AI had solved commonsense reasoning. Instead, the WSC became an important example of a broader problem in AI evaluation: a system can master a benchmark without necessarily mastering the general capability that the benchmark was intended to measure.

Research published in 2024 reinforced this point. Prompted large language models performed strongly on Winograd-style datasets while performing substantially worse on other forms of pronoun ambiguity. This suggests that performance on any single benchmark provides only a partial picture of a model's underlying reasoning ability. "Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution" | Ian Porada and Jackie C. K. Cheung, 2024

The WSC therefore remains significant today, but primarily as a historical milestone in AI development and evaluation. It helped focus research on commonsense reasoning, contributed to later benchmarks such as SuperGLUE and WinoGrande, and demonstrated an important lesson for modern AI: benchmark success and general intelligence are not the same thing.

Earlier research also illustrates how approaches to the problem evolved. Before large pretrained language models became dominant, researchers explored explicit knowledge representation and symbolic reasoning. One approach used graph representations and Answer Set Programming (ASP) to incorporate additional commonsense knowledge and constraints, handling 240 of 291 WSC problems examined in that work. "Using Answer Set Programming for Commonsense Reasoning in the Winograd Schema Challenge" | Arpit Sharma

The Sentences Computers Can't Understand, But Humans Can

A clear historical introduction to Winograd schemas and why they were once considered such a difficult test of machine language understanding. The video's claim that computers couldn't solve them reflects the state of AI at the time; modern language models have since achieved strong performance on the benchmark.

Using ASP for Commonsense Reasoning in the Winograd Schema Challenge

A presentation of the ICLP 2019 paper showing a symbolic approach to the WSC using Answer Set Programming. It provides useful historical perspective on how researchers attempted to explicitly represent commonsense knowledge before large pretrained language models transformed the field.