Difference between revisions of "Moonshots"

From
Jump to: navigation, search
m (Agents & Agentic Workflows)
 
(24 intermediate revisions by the same user not shown)
Line 1: Line 1:
 
 
{{#seo:
 
{{#seo:
 
|title=PRIMO.ai
 
|title=PRIMO.ai
 
|titlemode=append
 
|titlemode=append
|keywords=ChatGPT, artificial, intelligence, machine, learning, NLP, NLG, NLC, NLU, models, data, singularity, moonshot, Sentience, AGI, Emergence, Moonshot, Explainable, TensorFlow, Google, Nvidia, Microsoft, Azure, Amazon, AWS, Hugging Face, OpenAI, Tensorflow, OpenAI, Google, Nvidia, Microsoft, Azure, Amazon, AWS, Meta, LLM, metaverse, assistants, agents, digital twin, IoT, Transhumanism, Immersive Reality, Generative AI, Conversational AI, Perplexity, Bing, You, Bard, Ernie, prompt Engineering LangChain, Video/Image, Vision, End-to-End Speech, Synthesize Speech, Speech Recognition, Stanford, MIT
+
|keywords=ChatGPT, artificial, intelligence, machine, learning, NLP, NLG, NLC, NLU, models, data, singularity, moonshot, AGI, artificial general intelligence, world models, recursive self improvement, AI scientist, autonomous research, autonomous invention, biology, longevity, drug discovery, materials discovery, fusion energy, weather prediction, climate, robotics, embodied intelligence, lifelong learning, multi-agent, personal superintelligence, brain computer interface, BCI, compute, datacenters, energy, efficiency, neuromorphic, photonic computing, evaluation, benchmarks, ARC-AGI, interpretability, mechanistic interpretability, alignment, explainable AI, dual-use, biosecurity, cybersecurity, governance, export controls, open weights, labor, economics, Sentience, model welfare, Emergence, TensorFlow, Google, DeepMind, Nvidia, Microsoft, Azure, Amazon, AWS, Hugging Face, OpenAI, Meta, Anthropic, LLM, metaverse, assistants, agents, digital twin, IoT, Transhumanism, Immersive Reality, Generative AI, Conversational AI
|description=Helpful resources for your journey with artificial intelligence; videos, articles, techniques, courses, profiles, and tools
+
|description=Helpful resources exploring AI moonshots, AGI milestones, autonomous science, world models, robotics, compute and energy constraints, evaluation, interpretability, governance, personal superintelligence, and transformative applications of artificial intelligence
 
}}
 
}}
 +
 
[https://www.youtube.com/results?search_query=moonshot+moon+shot+ai YouTube search...]
 
[https://www.youtube.com/results?search_query=moonshot+moon+shot+ai YouTube search...]
 
[https://www.quora.com/search?q=ai%20moonshot%20moon ... Quora search]
 
[https://www.quora.com/search?q=ai%20moonshot%20moon ... Quora search]
Line 12: Line 12:
 
[https://www.bing.com/news/search?q=moonshot+moon+shot+ai&qft=interval%3d%228%22 ...Bing News]
 
[https://www.bing.com/news/search?q=moonshot+moon+shot+ai&qft=interval%3d%228%22 ...Bing News]
  
* [[Artificial General Intelligence (AGI) to Singularity]] ... [[Inside Out - Curious Optimistic Reasoning| Curious Reasoning]] ... [[Emergence]] ... [[Moonshots]] ... [[Explainable / Interpretable AI|Explainable AI]] ... [[Algorithm Administration#Automated Learning|Automated Learning]]
+
* [[Artificial General Intelligence (AGI) to Singularity]] ... [[Inside Out - Curious Optimistic Reasoning| Curious Reasoning]] ... [[Emergence]] ... [[Moonshots]] ... [[Explainable / Interpretable AI|Explainable AI]] ... [[Algorithm Administration#Automated Learning|Automated Learning]]
 +
* [[Agents]] ... [[Robotic Process Automation (RPA)|Robotic Process Automation]] ... [[Assistants]] ... [[Personal Companions]] ... [[Personal Productivity|Productivity]] ... [[Email]] ... [[Negotiation]] ... [[LangChain]]
 
* [[Large Language Model (LLM)]] ... [[Large Language Model (LLM)#Multimodal|Multimodal]] ... [[Foundation Models (FM)]] ... [[Generative Pre-trained Transformer (GPT)|Generative Pre-trained]] ... [[Transformer]] ... [[Attention]] ... [[Generative Adversarial Network (GAN)|GAN]] ... [[Bidirectional Encoder Representations from Transformers (BERT)|BERT]]
 
* [[Large Language Model (LLM)]] ... [[Large Language Model (LLM)#Multimodal|Multimodal]] ... [[Foundation Models (FM)]] ... [[Generative Pre-trained Transformer (GPT)|Generative Pre-trained]] ... [[Transformer]] ... [[Attention]] ... [[Generative Adversarial Network (GAN)|GAN]] ... [[Bidirectional Encoder Representations from Transformers (BERT)|BERT]]
 
* [[In-Context Learning (ICL)]] ... [[Context]] ... [[Out-of-Distribution (OOD) Generalization]]
 
* [[In-Context Learning (ICL)]] ... [[Context]] ... [[Out-of-Distribution (OOD) Generalization]]
* [[Immersive Reality]] ... [[Metaverse]] ... [[Omniverse]] ... [[Transhumanism]] ... [[Religion]]  
+
* [[Data Science]] ... [[Data Governance|Governance]] ... [[Data Preprocessing|Preprocessing]] ... [[Feature Exploration/Learning|Exploration]] ... [[Data Interoperability|Interoperability]] ... [[Algorithm Administration#Master Data Management (MDM)|Master Data Management (MDM)]] ... [[Bias and Variances]] ... [[Benchmarks]] ... [[Datasets]]  
* [[Telecommunications]] ... [[Computer Networks]] ... [[Telecommunications#5G|5G]] ... [[Satellite#Satellite Communications|Satellite Communications]] ... [[Quantum Communications]] ... [[Agents#Communication | Communication Agents]] ... [[Smart Cities]] ... [[Digital Twin]] ... [[Internet of Things (IoT)]]  
+
* [[Immersive Reality]] ... [[Metaverse]] ... [[Omniverse]] ... [[Transhumanism]] ... [[Religion]]
* [[Time]] ... [[Time#Positioning, Navigation and Timing (PNT)|PNT]] ... [[Time#Global Positioning System (GPS)|GPS]] ... [[Causation vs. Correlation#Retrocausality| Retrocausality]] ... [[Quantum#Delayed Choice Quantum Eraser|Delayed Choice Quantum Eraser]] ... [[Quantum]]
+
* [[Telecommunications]] ... [[Computer Networks]] ... [[Telecommunications#5G|5G]] ... [[Satellite#Satellite Communications|Satellite Communications]] ... [[Quantum Communications]] ... [[Agents#Communication | Communication Agents]] ... [[Smart Cities]] ... [[Digital Twin]] ... [[Internet of Things (IoT)]]
 
* [[Creatives]] ... [[History of Artificial Intelligence (AI)]] ... [[Neural Network#Neural Network History|Neural Network History]] ... [[Rewriting Past, Shape our Future]] ... [[Archaeology]] ... [[Paleontology]]
 
* [[Creatives]] ... [[History of Artificial Intelligence (AI)]] ... [[Neural Network#Neural Network History|Neural Network History]] ... [[Rewriting Past, Shape our Future]] ... [[Archaeology]] ... [[Paleontology]]
 
* [[Humor]]
 
* [[Humor]]
 
** [https://www.vice.com/en/article/z43nke/joke-telling-robots-are-the-final-frontier-of-artificial-intelligence Joke-Telling Robots Are the Final Frontier of Artificial Intelligence | Becky Ferreira - Vice] ... [[Humor]] requires self-awareness, spontaneity, linguistic sophistication, and empathy. Not easy for a [[Robotics | robot]].
 
** [https://www.vice.com/en/article/z43nke/joke-telling-robots-are-the-final-frontier-of-artificial-intelligence Joke-Telling Robots Are the Final Frontier of Artificial Intelligence | Becky Ferreira - Vice] ... [[Humor]] requires self-awareness, spontaneity, linguistic sophistication, and empathy. Not easy for a [[Robotics | robot]].
*[https://medium.com/applied-data-science/how-to-build-your-own-alphazero-ai-using-python-and-keras-7f664945c188 How to build your own AlphaZero AI using Python and Keras]
 
 
* [[What is Artificial Intelligence (AI)? | Artificial Intelligence (AI)]] ... [[Generative AI]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]]
 
* [[What is Artificial Intelligence (AI)? | Artificial Intelligence (AI)]] ... [[Generative AI]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]]
 
* [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[Bing/Copilot]] | [[Microsoft]] ... [[Gemini]] | [[Google]] ... [[Claude]] | [[Anthropic]] ... [[Perplexity]] ... [[You]] ... [[phind]] ... [[Ernie]] | [[Baidu]]
 
* [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[Bing/Copilot]] | [[Microsoft]] ... [[Gemini]] | [[Google]] ... [[Claude]] | [[Anthropic]] ... [[Perplexity]] ... [[You]] ... [[phind]] ... [[Ernie]] | [[Baidu]]
* [https://www.nature.com/articles/s41586-024-07936-z The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Sakana AI Team - Nature] ...Breakthrough in AI-generated research papers achieving human-level acceptance at workshops.
+
* [[Data Science]] ... [[Data Governance|Governance]] ... [[Data Preprocessing|Preprocessing]] ... [[Feature Exploration/Learning|Exploration]] ... [[Data Interoperability|Interoperability]] ... [[Algorithm Administration#Master Data Management (MDM)|Master Data Management (MDM)]] ... [[Bias and Variances]] ... [[Benchmarks]] ... [[Datasets]]
* [https://www.congress.gov/bill/119th-congress/senate-bill/AI-Grand-Challenges-Act AI Grand Challenges Act | U.S. Senate - 2026] ...Proposed bipartisan legislation to fund prize competitions for AI breakthroughs in health and security.
+
* [[Risk, Compliance and Regulation]]  ... [[Ethics]]  ... [[Privacy]]  ... [[Law]]  ... [[AI Governance]]  ... [[AI Verification and Validation]]
 +
* [https://medium.com/applied-data-science/how-to-build-your-own-alphazero-ai-using-python-and-keras-7f664945c188 How to build your own AlphaZero AI using Python and Keras]
 +
* [https://www.nature.com/articles/s41586-026-10265-5 The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Sakana AI, UBC, Vector Institute, Oxford - Nature] ...Peer-reviewed account of an end-to-end automated research pipeline.
 +
* [https://www.pnas.org/doi/10.1073/pnas.2524472123 Large Language Models Pass a Standard Three-Party Turing Test | Cameron Jones & Benjamin Bergen - PNAS] ...The 76-year benchmark falls, and is immediately reinterpreted.
 +
* [https://www.booker.senate.gov/imo/media/doc/one_pager_bipartisan_ai_grand_challenges_act_of_2026.pdf Bipartisan AI Grand Challenges Act of 2026 | U.S. Senate] ...Directs NSF to run prize competitions with $1M minimums, and at least $10M for cancer breakthroughs.
 +
* [https://arxiv.org/pdf/2602.21012 International AI Safety Report 2026] ...Expert synthesis on capabilities, risks, and the disputed pace of research automation.
 +
 
 +
The “Sputnik” moment for [[Government Services#China|China]] came a year ago when a Google computer program, AlphaGo, beat the world’s top master of the ancient board game of Go. Now, [[Government Services#China|China]] is racing to become the world leader in artificial-intelligence. The term "moonshot"  is derived from the Apollo program, which was a series of space missions undertaken by the United States in the 1960s and early 1970s with the goal of landing humans on the Moon. The Apollo program was considered a moonshot because it represented a major technological and engineering challenge that required significant innovation and investment. In [[context]], what do you think would be a "Moonshot" response?
  
The “Sputnik” moment for [[Government Services#China|China]] came a year ago when a Google computer program, AlphaGo, beat the world’s top master of the ancient board game of Go. Now, [[Government Services#China|China]] is racing to become the world leader in artificial-intelligence. In [[context]], what do you think would be a "Moonshot" response?
 
  
 
<hr><center>
 
<hr><center>
  
<b><i>The 'moonshot' milestones along the road to Artificial General Intelligence (AGI)</i></b>
+
<b><i>The "moonshot" milestones along the road to Artificial General Intelligence (AGI)</i></b>
  
 
</center><hr>
 
</center><hr>
  
In the [[context]] of AI, a "moonshot" refers to a project or goal that aims to achieve a major breakthrough in artificial intelligence that has the potential to transform society or address significant global challenges. The term "moonshot" is derived from the Apollo program, which was a series of space missions undertaken by the United States in the 1960s and early 1970s with the goal of landing humans on the Moon. The Apollo program was considered a moonshot because it represented a major technological and engineering challenge that required significant innovation and investment.
+
In the [[context]] of AI, a "moonshot" refers to a project or goal that aims to achieve a major breakthrough in artificial intelligence that has the potential to transform society or address significant global challenges. Moonshots are high-risk, high-impact goals where advanced AI is used to tackle problems that previously seemed decades away—such as autonomous scientific discovery, general-purpose robotics, radical health extension, new energy technologies, and eventually AI systems that improve AI itself. In short, a moonshot is a high-risk, high-impact goal that aims to achieve a major breakthrough in machine intelligence or use advanced AI to address problems that previously appeared decades away.
 +
 
 +
AI moonshots fall into two overlapping categories:
 +
 
 +
* '''Capability Moonshots''' — creating AI systems that can understand the world, reason, learn continuously, conduct research, operate autonomously, improve other AI systems, and act effectively in the physical world.
 +
* '''Civilization-Scale Moonshots''' — applying those capabilities to scientific discovery, medicine, longevity, energy, materials, climate, education, and other major human challenges.
 +
 
 +
The frontier therefore extends beyond simply making a more capable chatbot. The deeper question is:
 +
 
 +
:'''''What becomes possible when machine intelligence can understand, discover, invent, learn, act, and improve at increasingly general levels?'''''
 +
 
 +
And immediately behind it, a second question that the first cannot answer on its own:
 +
 
 +
:'''''What has to be true — physically, economically, and institutionally — for that to happen safely?'''''
 +
 
 +
= Calibration: Moonshots Already Achieved =
 +
 
 +
Before surveying what remains, it is worth recording what has already fallen. Each of these was, in its time, described as a definitive test of machine intelligence. Each was reinterpreted as "not real intelligence" shortly after it was accomplished — a pattern sometimes called the '''AI effect'''.
 +
 
 +
{| class="wikitable"
 +
! Milestone !! Roughly When !! How It Was Reinterpreted
 +
|-
 +
| Checkers solved || 1990s–2007 || Search, not thought
 +
|-
 +
| Chess grandmaster defeat || 1997 || Brute force, not understanding
 +
|-
 +
| Jeopardy! championship || 2011 || Retrieval, not comprehension
 +
|-
 +
| Go world champion defeat || 2016 || Narrow domain with clean rules
 +
|-
 +
| Protein structure prediction || 2020–2024 || Pattern-matching on a solved-enough problem
 +
|-
 +
| Competition-level mathematics (IMO gold) || 2025 || Exercises with known answers
 +
|-
 +
| Three-party Turing test || 2025–2026 || Imitation and substitutability, not intelligence
 +
|}
 +
 
 +
The Turing test case is instructive. In a peer-reviewed [https://www.pnas.org/doi/10.1073/pnas.2524472123 PNAS study], GPT-4.5 given a persona prompt was judged human 73% of the time — more often than the actual humans it was paired against — while a 1960s ELIZA control scored 23%, confirming that interrogators were not simply guessing. Yet the same models did '''not''' robustly pass without the persona prompt, and in some conditions performed no better than ELIZA. The authors frame the test as a measure of '''substitutability''' rather than intelligence, and warn that the near-term consequence is "counterfeit people" — erosion of online trust and displacement of genuine human contact.
 +
 
 +
Two lessons for everything below:
 +
 
 +
* '''Benchmarks fall faster than expected, and mean less than advertised.''' Plan for both.
 +
* '''Prompting, scaffolding, and tooling often carry more of a result than the underlying model.''' Always ask what was wrapped around the system before the milestone was declared.
 +
 
 +
= The Moonshot Index =
 +
 
 +
{| class="wikitable"
 +
! # !! Moonshot !! Category !! Core Question
 +
|-
 +
| 1 || World Models || Capability || Can AI simulate futures before acting?
 +
|-
 +
| 2 || Recursive Self-Improvement || Capability || Can AI improve AI?
 +
|-
 +
| 3 || Autonomous Research & Discovery || Capability || Can AI run the whole discovery cycle?
 +
|-
 +
| 4 || Autonomous Invention & Engineering || Capability || Can AI build what we haven't?
 +
|-
 +
| 5 || Unsolved Mathematics & Formal Verification || Capability || Can AI discover new mathematics?
 +
|-
 +
| 6 || Embodied General Intelligence || Capability || Can AI act in the messy physical world?
 +
|-
 +
| 7 || Lifelong Learning || Capability || Can AI keep learning after training?
 +
|-
 +
| 8 || Multi-Agent Ecosystems || Capability || What happens when agents deal with agents?
 +
|-
 +
| 9 || Radical Efficiency || Capability || Can intelligence run on 20 watts?
 +
|-
 +
| 10 || Biology, Medicine & Longevity || Civilization || Can AI repair and redesign living systems?
 +
|-
 +
| 11 || Materials & Energy || Civilization || Can AI find technologies we never found?
 +
|-
 +
| 12 || Weather, Climate & Earth Systems || Civilization || Can AI forecast and manage a planet?
 +
|-
 +
| 13 || Education || Civilization || A world-class tutor for every learner?
 +
|-
 +
| 14 || Personal Superintelligence || Civilization || Superintelligence for individuals, not just institutions?
 +
|-
 +
| 15 || Space & Astronomy || Civilization || Can AI explore where we cannot go?
 +
|-
 +
| 16 || Decoding Non-Human Communication || Civilization || Can AI let us talk to other species?
 +
|-
 +
| 17 || Understanding Intelligence Itself || Civilization || Can AI explain the brain that built it?
 +
|-
 +
| 18 || Compute, Energy & Hardware Substrate || Enabling || Can we build the launch pad?
 +
|-
 +
| 19 || Evaluation & Measurement || Enabling || How would we know we succeeded?
 +
|-
 +
| 20 || Interpretability || Guardrail || Can we see inside the systems we build?
 +
|-
 +
| 21 || Reliability, Alignment & Control || Guardrail || Can powerful AI stay correctable?
 +
|-
 +
| 22 || Misuse, Dual-Use & Security || Guardrail || Can capability be released without arming the wrong hands?
 +
|-
 +
| 23 || Governance, Geopolitics & Access || Guardrail || Who decides, and who benefits?
 +
|-
 +
| 24 || Economics & Labor || Guardrail || What happens to work?
 +
|-
 +
| 25 || Moral Status of AI || Open || Does any of this matter to the systems themselves?
 +
|}
 +
 
 +
= Part I — Capability Moonshots =
 +
 
 +
== 1. World Models — AI That Can Simulate Possible Futures ==
 +
 
 +
[https://www.youtube.com/results?search_query=AI+world+models+Genie+3+DeepMind+AGI YouTube search...]
 +
[https://www.google.com/search?q=AI+world+models+Genie+3+DeepMind+AGI ...Google search]
  
== Able to Predict the Future ==
+
Humans don't merely react to the world. We continuously construct internal models of it — anticipating what might happen next, imagining alternatives, and considering the consequences of our actions before we act.
[https://www.youtube.com/results?search_query=predict+future+artificial+intelligence Youtube search...]
 
[https://www.google.com/search?q=predict+future+artificial+intelligence ...Google search]
 
  
<youtube>Wf1UFz2jAJU</youtube>
+
'''World models''' attempt to give AI a similar capability.
<youtube>VrIMLt__CJc</youtube>
 
<youtube>5olCp-JEzrI</youtube>
 
<youtube>8z0Mrbj34uY</youtube>
 
  
== Can Conjure & Ask Questions ==
+
Rather than only recognizing an image, generating text, or responding to the present situation, a world model can represent how an environment changes through time and predict how different actions could alter what happens next.
[https://www.youtube.com/results?search_query=ai+ask+question+cognitive+computing+general+artificial+intelligence Youtube search...]
 
[https://www.google.com/search?q=ai+ask+question+cognitive+computing+general+artificial+intelligence ...Google search]
 
  
* [[Inside Out - Curious Optimistic Reasoning]]
+
The progression is:
* [https://www.youtube.com/results?search_query=smartest+animals Animals]
 
* [https://www.youtube.com/results?search_query=general+artificial+intelligence General Intelligence]
 
* [https://www.youtube.com/results?search_query=consciousness+artificial+intelligence Consciousness]
 
  
<youtube>gtThWYWOyFQ</youtube>
+
:'''''Perception → World Model → Imagine Futures → Choose Action'''''
<youtube>L4QYUEz0mI8</youtube>
 
  
== Embodied General Intelligence ==
+
[[Google DeepMind]] describes its Genie 3 system as a general-purpose world model capable of generating interactive environments in real time. These environments can be used by AI agents to learn how situations evolve and how their own actions affect the environment. DeepMind considers world models an important step toward [[Artificial General Intelligence (AGI)|AGI]]. [https://deepmind.google/models/genie/ "Genie 3 — A New Frontier for World Models" | Google DeepMind]
Many people call this the "physical Turing test." Embodied General Intelligence is all about getting robots to navigate messy, unpredictable real-world spaces. Instead of repeating the same motion on an assembly line, these robots need to figure out how to fold laundry or do the dishes in a kitchen they have never seen before.
+
 
 +
World models could eventually allow AI systems to rehearse actions internally before attempting them in the real world — similar to imagining several possible moves before making a decision.
 +
 
 +
This capability is especially important for [[Agents]], [[Robotics]], autonomous vehicles, planning, scientific simulation, and long-horizon decision making.
 +
 
 +
It also makes possible a deeper form of reasoning:
 +
 
 +
:'''''What will happen? → What could happen? → What happens if I intervene? → Which future should I try to create?'''''
 +
 
 +
'''Open problems:''' consistency over long horizons, physical accuracy versus visual plausibility, and whether a model that generates convincing video has actually learned physics or merely learned what physics looks like.
 +
 
 +
<youtube>PAU2VeoVplg</youtube>
 +
 
 +
== 2. Recursive Self-Improvement ==
 +
 
 +
[https://www.youtube.com/results?search_query=recursive+self+improvement+AI+automated+AI+research YouTube search...]
 +
[https://www.google.com/search?q=recursive+self+improvement+AI+automated+AI+research ...Google search]
 +
 
 +
The topic of '''recursive self-improvement''' is a significant threshold in AI development. Currently, much of the process is still driven by human software engineers and researchers who generate training data, run ablations to test data quality, design experiments, and evaluate models against benchmarks.
 +
 
 +
However, many laboratories are working to close more of this loop through automation:
 +
 
 +
* '''Automated Judges:''' Different models can evaluate the quality and correctness of outputs.
 +
* '''Generative Feedback:''' Models can generate new, high-quality training examples and feedback.
 +
* '''Adversarial Reasoning:''' Advanced models can compare, criticize, test, and filter candidate solutions.
 +
* '''Automated Experimentation:''' Agents can modify code, run experiments, evaluate results, and propose subsequent experiments.
 +
 
 +
The resulting cycle increasingly resembles:
 +
 
 +
:'''''Generate → Evaluate → Select → Train → Test → Improve → Repeat'''''
 +
 
 +
By feeding successful results back into AI research and post-training, development speed could increase substantially.
 +
 
 +
More advanced forms of this process could eventually allow AI systems to contribute directly to the design of their successors.
 +
 
 +
:'''''AI helps improve AI → improved AI becomes better at improving AI → the research cycle accelerates.'''''
 +
 
 +
This possibility is sometimes associated with an '''intelligence explosion'''. Mustafa Suleyman and others caution that such a process would still face substantial limits involving compute, experimentation, infrastructure, reliability, and human control.
 +
 
 +
'''How much acceleration? Experts disagree sharply.''' In a forecasting study summarized in the [https://arxiv.org/pdf/2602.21012 International AI Safety Report 2026], AI-specialist forecasters gave a median 20% probability that the next few years could compress six years of prior advancement into two, while generalist superforecasters put it at 8%. Estimates rose for scenarios in which AI systems outperform humans on month-long research projects. Empirical evidence on whether research automation is actually accelerating progress remains mixed. See also [[#19. Evaluation & Measurement — How Would We Know?|Moonshot 19]], since this claim is unusually hard to measure from outside a frontier lab.
 +
 
 +
<youtube>-RXD4bTuFTo</youtube> <youtube>tQ5wO1lznCQ</youtube>
 +
 
 +
== 3. Autonomous AI Research & Discovery ==
 +
 
 +
The frontier of AI is shifting from static chat interfaces toward increasingly autonomous '''agentic workflows'''. Rather than performing only one prompt-response cycle, agents can plan tasks, use tools, write and execute code, evaluate results, revise their approach, and coordinate with other agents.
 +
 
 +
Scientific research is becoming one of the most important proving grounds for these capabilities.
 +
 
 +
The larger moonshot is:
 +
 
 +
:'''''Can AI participate in the entire discovery cycle — from asking the question to proposing the hypothesis, designing the experiment, interpreting the evidence, and deciding what to try next?'''''
 +
 
 +
=== 3a. The Automated Researcher ===
 +
 
 +
By September 2026, OpenAI reported hitting a self-set target: an '''automated research intern''', defined by the company as a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. The organization reported logging roughly 3.1 agent-workdays of effort per eight hours of human labor, with a growing share of researchers running four or more agents concurrently. The stated next target is a fully '''automated AI researcher''' by March 2028.
 +
 
 +
'''Read the definition carefully.''' Three qualifiers carry the weight: the tasks are ''well-defined'', a human is ''steering'', and the horizon is ''days'' rather than months. OpenAI set the goal, wrote the definition, took the measurements, and published the grade; no external party can independently verify it. The company itself notes that hard-to-automate tasks and compute availability could constrain further progress. Treat it as a credible internal measurement of a self-authored target, not as an externally validated milestone.
 +
 
 +
<youtube>pCoQtk0qHGk</youtube>
 +
 
 +
=== 3b. AI as Scientist ===
 +
 
 +
We have moved past seeing AI only as a tool for retrieving scientific information. Increasingly capable systems can help generate hypotheses, review scientific literature, design experiments, analyze results, and critique possible explanations.
 +
 
 +
[[Google DeepMind]]'s Co-Scientist, for example, uses multiple Gemini-based agents to generate, debate, criticize, and refine scientific hypotheses. [https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/ "Co-Scientist: A Multi-Agent AI Partner to Accelerate Research" | Google DeepMind]
 +
 
 +
Sakana AI's "AI Scientist" explores another version of the concept: systems capable of generating research ideas, running computational experiments, analyzing results, and producing scientific papers. The work was published in [https://www.nature.com/articles/s41586-026-10265-5 Nature], and an unedited, fully AI-generated manuscript passed blind human peer review at an ICLR workshop, scoring above the average human acceptance threshold.
 +
 
 +
The transition is:
 +
 
 +
:'''''AI answers scientific questions → AI helps investigate scientific questions → AI proposes new scientific questions.'''''
 +
 
 +
'''The honest caveat:''' most such systems are still evaluated on paper generation and benchmark optimization rather than on discoveries that change a field. Verification — knowing whether an AI-generated result is actually correct — is emerging as the binding constraint, and has spawned its own benchmark literature.
 +
 
 +
<youtube>RnISO02IBcw</youtube>
 +
 
 +
== 4. Autonomous Invention & Engineering ==
 +
 
 +
Scientific discovery asks '''what is true?'''
 +
 
 +
Engineering asks '''what can we build?'''
 +
 
 +
A further AI moonshot is therefore autonomous invention — systems capable of designing algorithms, software, chips, machines, materials, and other technologies that human engineers have not previously created.
 +
 
 +
Google DeepMind's [[AlphaEvolve]] combines large language models with automated evaluators and an evolutionary search process to discover and optimize algorithms. Its discoveries have been applied to computing infrastructure, chip design, AI training, mathematics, quantum computing, and other technical problems. [https://deepmind.google/blog/alphaevolve-impact/ "AlphaEvolve: How Our Gemini-Powered Coding Agent Is Scaling Impact Across Fields" | Google DeepMind]
 +
 
 +
Similarly, Gemini Deep Think is being explored for real-world engineering problems including semiconductor fabrication and mechanical design. [https://deepmind.google/models/gemini/deep-think/ "Gemini Deep Think" | Google DeepMind]
 +
 
 +
The progression is:
 +
 
 +
:'''''Solve our problems → Discover new principles → Invent new technologies'''''
 +
 
 +
Eventually, AI-assisted engineering could create technologies whose designs are too complex for any individual human to derive unaided. Note the loop this closes with [[#18. Compute, Energy & Hardware Substrate|Moonshot 18]]: AI that designs better chips is AI that improves its own substrate.
 +
 
 +
== 5. Unsolved Mathematics & Formal Verification ==
 +
 
 +
After reaching gold-medal-level performance on International Mathematical Olympiad problems, frontier AI systems are increasingly being tested on original mathematical research.
 +
 
 +
The deeper moonshot isn't simply solving difficult exercises whose answers are already known.
 +
 
 +
It is:
 +
 
 +
:'''''Can AI discover genuinely new mathematics?'''''
 +
 
 +
Systems can increasingly combine neural reasoning with formal verification tools such as [[Lean]], allowing a proposed proof to be checked step-by-step against rigorous logical rules.
 +
 
 +
Formal verification acts like a highly sophisticated proof checker: instead of merely judging whether an argument sounds plausible, the system verifies that each logical step follows correctly.
 +
 
 +
This combination of creative reasoning and machine-verifiable proof could make mathematics one of the first domains in which highly autonomous AI research becomes possible — precisely because the verification problem that limits [[#3b. AI as Scientist|Moonshot 3b]] is, uniquely here, solved.
 +
 
 +
Concrete targets include long-standing open problems, large-scale formalization of existing mathematical literature, and AI-discovered proofs that human mathematicians accept as novel rather than rediscovered.
 +
 
 +
<youtube>-aXkChrQl1I</youtube>
 +
 
 +
== 6. Embodied General Intelligence ==
 +
 
 +
[https://www.youtube.com/results?search_query=embodied+AI+general+intelligence+robotics+physical+AI YouTube search...]
 +
[https://www.google.com/search?q=embodied+AI+general+intelligence+robotics+physical+AI ...Google search]
 +
 
 +
Many researchers describe this as a kind of '''physical Turing test''': a robot that can walk into an unfamiliar home and do the laundry.
 +
 
 +
Embodied General Intelligence is about moving intelligence from the digital world into messy, unpredictable physical environments.
 +
 
 +
Traditional industrial robots repeat carefully programmed movements in controlled environments. A generally capable robot must instead perceive unfamiliar surroundings, understand objects, predict physical consequences, recover from errors, and determine how to accomplish goals it has never encountered before.
 +
 
 +
A household robot, for example, can't simply memorize one sequence for folding laundry or loading a dishwasher. Every home, object, obstruction, person, and situation will be different.
 +
 
 +
The progression is:
 +
 
 +
:'''''See the World → Understand the World → Act in the World → Adapt to the Unexpected'''''
 +
 
 +
World models are closely connected to this challenge. Robots may eventually learn many physical skills inside simulated environments before transferring those skills to real machines.
 +
 
 +
=== 6a. The Data Bottleneck ===
 +
 
 +
Language models were trained on an internet-scale corpus that already existed. '''No equivalent corpus exists for physical manipulation.''' This is the central reason robotics lags language, and it defines the sub-moonshots:
 +
 
 +
* '''Teleoperation at scale''' — humans puppeting robots to generate demonstration data, which is slow and expensive.
 +
* '''Sim2real transfer''' — training in simulation where data is free, then closing the reality gap.
 +
* '''Learning from human video''' — extracting manipulation knowledge from ordinary footage of people doing things.
 +
* '''Robot foundation models''' — general-purpose policies transferable across different robot bodies rather than one model per machine.
 +
* '''Shared fleet learning''' — every robot's experience improving every other robot.
 +
 
 +
:'''''No data → borrowed data → simulated data → fleet-generated data → self-sustaining physical learning'''''
  
 
<youtube>lK7TjujKQLw</youtube>
 
<youtube>lK7TjujKQLw</youtube>
  
=== Autonomous Vehicles ===
+
=== 6b. Autonomous Vehicles ===
[https://www.youtube.com/results?search_query=Autonomous+vehicles+self+driving+autocomplete+artificial+intelligence Youtube search...]
+
 
[https://www.google.com/search?q=Autonomous+vehicles+self+driving+autocomplete+artificial+intelligence ...Google search]
+
[https://www.youtube.com/results?search_query=Autonomous+vehicles+self+driving+artificial+intelligence YouTube search...]
 +
[https://www.google.com/search?q=Autonomous+vehicles+self+driving+artificial+intelligence ...Google search]
 +
 
 
* [[Transportation (Autonomous Vehicles)]]
 
* [[Transportation (Autonomous Vehicles)]]
 +
 +
Autonomous vehicles represent an important early example of embodied intelligence operating in the open world.
 +
 +
Driving requires continuous perception, prediction, planning, spatial reasoning, risk assessment, and interaction with unpredictable human behavior.
 +
 +
A successful autonomous vehicle must constantly answer:
 +
 +
:'''''What is happening? → What is likely to happen next? → What should I do about it?'''''
 +
 +
Autonomous driving is also the field's most instructive lesson in the gap between demonstration and deployment: impressive prototypes arrived more than a decade before broad commercial operation, and the delay was almost entirely about the last few percent of reliability. Expect the same shape for household robotics and autonomous research.
  
 
<youtube>yt015gM-ync</youtube>
 
<youtube>yt015gM-ync</youtube>
  
== Recursive Self Improvement ==
+
== 7. Lifelong Learning — AI That Keeps Learning ==
The topic of recursive self-improvement is a significant threshold in AI development. Currently, the process is largely driven by human software engineers who manually generate training data, run ablations to test data quality, and evaluate models against benchmarks.
+
 
 +
[https://www.youtube.com/results?search_query=lifelong+learning+continual+learning+AI+agents+memory YouTube search...]
 +
[https://www.google.com/search?q=lifelong+learning+continual+learning+AI+agents+memory ...Google search]
 +
 
 +
Most modern AI systems undergo an enormous training process and are then largely fixed until their developers train or release another version.
 +
 
 +
Human intelligence works differently.
 +
 
 +
We continuously accumulate experiences, form memories, acquire skills, revise beliefs, adapt to changing environments, and learn from both success and failure.
 +
 
 +
A genuinely general intelligence may eventually require a similar ability to learn throughout its operational lifetime.
 +
 
 +
The progression is:
 +
 
 +
:'''''Static Model → Persistent Memory → Continual Learning → Lifelong Adaptive Intelligence'''''
 +
 
 +
Such a system could:
 +
 
 +
* Remember important experiences over months or years.
 +
* Learn new skills without requiring complete retraining.
 +
* Adapt to an individual user or environment.
 +
* Incorporate new knowledge while preserving older knowledge.
 +
* Learn from its own successes and failures.
 +
* Transfer lessons learned in one domain into another.
 +
 
 +
This sounds straightforward, but it creates difficult technical and safety problems.
 +
 
 +
New learning must not erase important existing capabilities — a problem associated with '''catastrophic forgetting'''. An adaptive system must also avoid gradually drifting away from its intended goals, learning harmful behaviors, incorporating false information, or modifying itself in ways humans no longer understand.
 +
 
 +
Lifelong learning therefore connects directly to [[Memory]], [[Agents]], recursive self-improvement, personal superintelligence, and [[Explainable / Interpretable AI|AI alignment]].
 +
 
 +
The moonshot is:
 +
 
 +
:'''''An intelligence that doesn't simply arrive trained — it continues growing through experience.'''''
 +
 
 +
== 8. Multi-Agent Ecosystems — When Agents Deal With Agents ==
 +
 
 +
[https://www.youtube.com/results?search_query=multi+agent+AI+ecosystems+agent+protocols+agentic+commerce YouTube search...]
 +
[https://www.google.com/search?q=multi+agent+AI+ecosystems+agent+protocols+agentic+commerce ...Google search]
 +
 
 +
Most discussion of agents imagines one agent serving one person. The more consequential scenario is '''many agents interacting with each other''', often without a human in the loop for any individual exchange.
 +
 
 +
This is already emerging: agents negotiating with agents, agents transacting on behalf of users, agents delegating subtasks to specialist agents, and interoperability protocols that let systems from different vendors discover and call one another.
 +
 
 +
:'''''One agent, one task → agent teams → agent markets → machine-to-machine economies'''''
 +
 
 +
The moonshot is a functioning agent ecosystem with:
 +
 
 +
* '''Identity and authentication''' — knowing which agent acts for whom.
 +
* '''Delegation and authority limits''' — bounded permission to spend, commit, or act.
 +
* '''Payment and settlement rails''' designed for machine counterparties.
 +
* '''Dispute resolution and liability''' when an agent errs on someone's behalf.
 +
* '''Interoperability standards''' so agents can discover and use each other's tools.
 +
 
 +
The risks are distinctive and poorly studied: collusion between agents, cascading failures, flash-crash dynamics at machine speed, and emergent behavior no individual designer intended. Multi-agent dynamics are a live gap in the alignment literature, which has focused mostly on single systems.
 +
 
 +
== 9. Radical Efficiency — Intelligence on 20 Watts ==
 +
 
 +
[https://www.youtube.com/results?search_query=AI+efficiency+small+models+on+device+distillation+energy+per+token YouTube search...]
 +
[https://www.google.com/search?q=AI+efficiency+small+models+on+device+distillation+energy+per+token ...Google search]
 +
 
 +
The human brain performs general reasoning on roughly '''20 watts'''. Frontier AI systems require datacenters. That gap is not a footnote — it is one of the clearest measures of how far current approaches are from the thing they are imitating.
 +
 
 +
Efficiency is often treated as an engineering detail rather than a moonshot. It is a moonshot, and possibly the one with the shortest path to broad human impact, because it determines:
 +
 
 +
* Whether capable AI runs on a phone, a hearing aid, or a robot rather than in a distant datacenter.
 +
* Whether [[#14. Personal Superintelligence — Powerful AI for Every Individual|personal superintelligence]] is affordable for billions of people or a subscription for the wealthy.
 +
* Whether inference cost stops constraining how much reasoning a system is allowed to do.
 +
* Whether AI's energy footprint is politically and environmentally sustainable.
 +
 
 +
:'''''Frontier capability in a datacenter → in a rack → on a laptop → on a phone → on 20 watts'''''
 +
 
 +
The competitive dimension is real. Chinese laboratories have pushed hard on the cost-per-capability frontier: Moonshot AI's Kimi K3, at roughly 2.8 trillion parameters, was released as an open-weight model claimed to match or exceed leading Western systems on mainstream benchmarks at a fraction of the running cost — a release that moved markets and prompted a reassessment of assumed US leads. Efficiency, not raw scale, is increasingly where the race is run.
 +
 
 +
= Part II — Civilization-Scale Moonshots =
 +
 
 +
== 10. AI for Biology, Medicine & Longevity — Understand and Engineer Life ==
 +
 
 +
[https://www.youtube.com/results?search_query=AI+biology+medicine+longevity+AlphaFold+Rosalind+AI+scientist YouTube search...]
 +
[https://www.google.com/search?q=AI+biology+medicine+longevity+AlphaFold+Rosalind+AI+scientist ...Google search]
 +
 
 +
Biology is extraordinarily complex. A living organism contains interacting systems spanning molecules, proteins, genes, cells, organs, environments, and behavior.
 +
 
 +
AI offers the possibility of modeling these systems at scales that would be extremely difficult for humans to analyze unaided.
 +
 
 +
[[AlphaFold]] demonstrated the potential by transforming protein-structure prediction. The larger moonshot is to move progressively from '''describing biology''' toward '''predicting and eventually designing biological outcomes'''.
 +
 
 +
The progression could be:
 +
 
 +
:'''''Understand Biology → Predict Biology → Design Biology → Prevent or Reverse Disease'''''
 +
 
 +
[[OpenAI]]'s GPT-Rosalind is a purpose-built frontier reasoning model for life-sciences research, covering molecules, proteins, genes, pathways, and disease biology, with emphasis on multi-step tool use across literature review, sequence-to-function interpretation, experimental planning, and data analysis. It is deployed as a research preview through a trusted-access program with biopharma and research partners, and is evaluated against an internally designed, externally expert-judged benchmark spanning six workflow areas of life-sciences research. [https://openai.com/index/introducing-gpt-rosalind/ "Introducing GPT-Rosalind for Life Sciences Research" | OpenAI]
 +
 
 +
Google DeepMind's AI Co-Scientist similarly explores the use of multi-agent reasoning to generate and refine biological hypotheses. [https://deepmind.google/blog/co-scientist-a-multi-agent-ai-partner-to-accelerate-research/ "Co-Scientist" | Google DeepMind]
  
However, many labs are now working to close this loop to automate the process:
+
Possible long-term moonshots include:
  
* Automated Judges: Different models will act as judges to evaluate the quality of outputs.
+
* Designing new medicines rather than screening primarily from known compounds.
* Generative Feedback: Models will generate new, high-quality training data autonomously.
+
* Predicting how genetic changes affect cells and organisms.
* Adversarial Reasoning: Advanced models will reason over which data to include, effectively filtering for quality.
+
* Modeling complete cells — the '''virtual cell''' — and eventually larger biological systems.
 +
* Creating personalized treatments based on an individual's biology.
 +
* Detecting disease before symptoms appear.
 +
* Understanding the biological mechanisms of aging.
 +
* Extending healthy human lifespan.
  
By feeding this output back into the post-training process, development speed will likely increase significantly. While some speculate this could lead to an intelligence explosion, Mustafa Suleyman notes that achieving this requires substantial compute and, without proper human oversight or control, it introduces significant risks.
+
The rate limiter is rarely the idea. Target-to-approval for a new drug in the United States takes roughly 10–15 years, and gains at the earliest discovery stages compound downstream. AI that shortens hypothesis generation still runs into wet-lab throughput, clinical trial duration, and regulatory process — which is why closed-loop robotic laboratories ([[#11. AI for Materials & Energy — Discover Technologies Humans Haven't Found|Moonshot 11]]) matter as much as better models.
  
<youtube>-RXD4bTuFTo</youtube>
+
The ultimate question is profound:
  
 +
:'''''Can intelligence understand living systems deeply enough to deliberately repair, redesign, and preserve them?'''''
  
The frontier of AI has shifted from static chat interfaces to autonomous '''"Agentic Workflows."''' These systems are designed for recursive self-improvement and multi-agent orchestration, moving beyond simple prompt-response cycles.
+
'''This is also the page's sharpest dual-use case.''' The same capabilities that accelerate drug discovery lower barriers to engineered biological harm. Frontier laboratories have responded with gated access, biodefense programs, and trusted-deployment models rather than open release. See [[#22. Misuse, Dual-Use & Security|Moonshot 22]].
  
=== The Automated Researcher ===
+
<youtube>P_fHJIYENdI</youtube>
By September 2026, OpenAI hit a major milestone with its "automated research intern." Think of it as a tireless assistant that handles well-defined research tasks while a human steers the ship. For every eight hours a person puts in, the system completes about three days' worth of work. The ultimate goal is to have a fully autonomous AI researcher running by March 2028.
 
  
<youtube>pCoQtk0qHGk</youtube>
+
== 11. AI for Materials & Energy — Discover Technologies Humans Haven't Found ==
 +
 
 +
[https://www.youtube.com/results?search_query=AI+materials+discovery+fusion+energy+DeepMind+GNoME YouTube search...]
 +
[https://www.google.com/search?q=AI+materials+discovery+fusion+energy+DeepMind+GNoME ...Google search]
 +
 
 +
Human civilization has repeatedly been transformed by the discovery of new materials and new sources of energy — bronze, steel, concrete, semiconductors, petroleum, nuclear energy, photovoltaic materials, lithium-ion batteries, and many others.
 +
 
 +
Yet discovering useful materials remains an enormous search problem. The number of possible chemical structures is vastly larger than scientists can test experimentally.
 +
 
 +
AI can search these spaces computationally.
 +
 
 +
Google DeepMind's GNoME system predicted millions of previously unknown crystal structures, including hundreds of thousands of candidates predicted to be stable. Potential applications include batteries, electronics, superconductors, solar technologies, and many other fields. [https://deepmind.google/blog/millions-of-new-materials-discovered-with-deep-learning/ "Millions of New Materials Discovered with Deep Learning" | Google DeepMind]
 +
 
 +
The discovery process can increasingly connect AI with automated laboratories:
 +
 
 +
:'''''AI predicts → Robots synthesize → Instruments test → Results return to AI → AI proposes the next experiment'''''
 +
 
 +
This creates another self-improving scientific loop. It is worth noting the standing critique: a predicted stable structure is not a synthesized material, and the gap between computational candidates and validated, manufacturable substances remains large and contested.
 +
 
 +
Energy represents an equally ambitious target.
 +
 
 +
[[Google DeepMind]] has applied machine learning and reinforcement learning to the extraordinarily difficult problem of controlling plasma inside fusion reactors. In 2025, DeepMind and Commonwealth Fusion Systems announced a partnership using AI simulation, optimization, and real-time control to help advance practical fusion energy. [https://deepmind.google/blog/bringing-ai-to-the-next-generation-of-fusion-energy/ "Bringing AI to the Next Generation of Fusion Energy" | Google DeepMind]
 +
 
 +
The larger chain could be:
 +
 
 +
:'''''AI → New Materials → Better Batteries / Solar / Computing → Fusion → Abundant Clean Energy'''''
 +
 
 +
The moonshot isn't simply making today's technology more efficient.
 +
 
 +
It is using AI to discover '''entirely new technologies that humans haven't yet found''' — and, circularly, to power the datacenters that the rest of this page depends on.
 +
 
 +
<youtube>dgBLVm2L1P4</youtube>
 +
 
 +
== 12. Weather, Climate & Earth Systems ==
 +
 
 +
[https://www.youtube.com/results?search_query=AI+weather+prediction+GraphCast+GenCast+climate+modeling YouTube search...]
 +
[https://www.google.com/search?q=AI+weather+prediction+GraphCast+GenCast+climate+modeling ...Google search]
 +
 
 +
Machine-learning weather models are among the clearest '''already-realized''' wins in AI for science — producing forecasts competitive with or better than traditional numerical weather prediction, at a small fraction of the computational cost, in seconds rather than hours on supercomputers.
 +
 
 +
That success reframes a larger moonshot: modeling the Earth as a coupled system.
 +
 
 +
:'''''Forecast the weather → forecast extremes → model the climate → evaluate interventions'''''
 +
 
 +
Sub-goals include:
 +
 
 +
* '''Extreme-event prediction''' with enough lead time to evacuate and prepare.
 +
* '''Kilometer-scale regional climate projection''' usable for actual infrastructure planning.
 +
* '''Wildfire, flood, and drought forecasting''' integrated with response systems.
 +
* '''Grid optimization and renewable forecasting''' — matching intermittent supply to demand.
 +
* '''Carbon capture materials''' discovered through the pipeline in Moonshot 11.
 +
* '''Earth digital twins''' that let policymakers test interventions before committing to them.
 +
 
 +
This moonshot is unusual in having a genuinely near-term humanitarian payoff, an existing track record, and comparatively low dual-use risk. It is also where AI's own energy footprint becomes uncomfortable to ignore.
 +
 
 +
== 13. Education — A World-Class Tutor for Every Learner ==
 +
 
 +
[https://www.youtube.com/results?search_query=AI+tutor+personalized+learning+two+sigma+problem YouTube search...]
 +
[https://www.google.com/search?q=AI+tutor+personalized+learning+Bloom+two+sigma+problem ...Google search]
 +
 
 +
Benjamin Bloom's '''two-sigma problem''' observed that students receiving one-on-one tutoring performed roughly two standard deviations better than those in conventional classrooms — a result no scalable intervention had matched, because individual human tutors cannot be provided to every learner on Earth.
 +
 
 +
AI is the first technology with a plausible claim to that scale.
 +
 
 +
:'''''Content delivery → adaptive practice → responsive tutoring → a patient expert who knows this learner'''''
 +
 
 +
The moonshot requires more than a chatbot that answers homework questions:
 +
 
 +
* '''Modeling the learner''' — knowing what they actually understand versus what they can recite.
 +
* '''Productive struggle''' — withholding the answer when the struggle is the learning.
 +
* '''Motivation and persistence''', not just explanation.
 +
* '''Assessment integrity''' in a world where any assignment can be completed by machine.
 +
* '''Working in every language''', including those with little training data.
 +
* '''Teacher augmentation''' rather than replacement.
 +
 
 +
The failure mode is specific and already visible: a system that supplies answers efficiently can produce fluent students who have learned less than they appear to have. Measuring genuine learning gains — not engagement, not satisfaction — is the hard part, and connects directly to [[#19. Evaluation & Measurement — How Would We Know?|Moonshot 19]].
 +
 
 +
== 14. Personal Superintelligence — Powerful AI for Every Individual ==
 +
 
 +
[https://www.youtube.com/results?search_query=personal+superintelligence+AI+agents+brain+computer+interface+BCI YouTube search...]
 +
[https://www.google.com/search?q=personal+superintelligence+AI+agents+brain+computer+interface+BCI ...Google search]
 +
 
 +
'''Personal superintelligence''' describes a future in which extremely capable [[Artificial Intelligence|AI]] isn't concentrated only in governments, corporations, or research laboratories, but is available directly to individuals.
 +
 
 +
Instead of simply answering questions, a personal AI could understand a person's goals, preferences, history, surroundings, and ongoing activities; reason about complex problems; and take actions on that person's behalf.
 +
 
 +
[[Meta]] has made this idea a central part of its long-term AI strategy. Mark Zuckerberg describes the goal as giving everyone access to a personal superintelligence that can help people create, learn, communicate, pursue their interests, and accomplish things that would previously have required teams of specialists. [https://www.meta.com/superintelligence/ "Personal Superintelligence" | Meta]
 +
 
 +
The concept represents a shift:
 +
 
 +
:'''''AI as a tool → AI as an assistant → AI as an agent → AI as a personal superintelligence.'''''
 +
 
 +
A sufficiently capable personal AI might continuously work with an individual across many areas of life — researching questions, teaching new skills, creating software and media, organizing information, communicating with other systems, monitoring projects, and coordinating other specialized AI agents.
 +
 
 +
Combined with lifelong learning and persistent memory, such a system could increasingly adapt to the individual rather than requiring the individual to continually explain their context to the AI.
 +
 
 +
'''This moonshot depends on Moonshot 9.''' "Personal" superintelligence delivered from a datacenter at frontier inference cost is not personal in any meaningful sense — it is rented, metered, and revocable. Efficiency determines whether this section describes a democratizing technology or a subscription tier.
 +
 
 +
=== 14a. From Screens to Continuous AI Interfaces ===
 +
 
 +
For personal superintelligence to become genuinely personal, the interface between humans and AI may also change.
 +
 
 +
Today, most people communicate with AI by typing, speaking, or sharing images.
 +
 
 +
Meta is developing [[smart glasses]] as a more continuous interface. Cameras, microphones, displays, and other sensors can allow an AI to see some of what its user sees, hear what the user hears, understand the surrounding context, and provide assistance throughout the day. Meta has suggested that AI-enabled glasses could eventually become a major personal computing platform. [https://www.meta.com/superintelligence/ "Personal Superintelligence" | Meta]
 +
 
 +
A more radical possibility is the [[Brain–computer interface|brain-computer interface]] (BCI).
 +
 
 +
A BCI establishes a communication pathway between neural activity and a computer. Instead of translating every intention into movements of the hands, eyes, or voice, some information can potentially move directly between the nervous system and a digital device.
 +
 
 +
Current implantable BCIs are primarily being developed as '''medical technologies'''. [[Neuralink]], for example, is testing systems that allow people with severe paralysis to control computers, phones, robotic devices, and other equipment using neural activity. [https://neuralink.com/updates/two-years-of-telepathy/ "Two Years of Telepathy" | Neuralink]
 +
 
 +
The longer-term possibility is much broader:
 +
 
 +
:'''''Brain → AI → Digital World'''''
 +
 
 +
A high-bandwidth BCI combined with increasingly capable personal AI could eventually create a much more direct relationship between human intention and machine intelligence.
  
=== AI as Scientist ===
+
The AI might infer what a person wants to accomplish, retrieve information, operate digital systems, or communicate with other AI agents with less dependence on keyboards, screens, or spoken commands.
We have moved past seeing AI as just a tool; it is now a true collaborator. Models like Sakana AI's "AI Scientist" are actually producing fresh, peer-review-quality research. They are running experiments in areas like virtual cell modeling and designing new drugs from scratch. It is like having a digital post-doc working around the clock in the lab.
 
  
<youtube>RnISO02IBcw</youtube>
+
This does '''not''' mean that brain implants are required for personal superintelligence. Wearable devices, voice interfaces, augmented-reality glasses, and other technologies may provide much of this capability without surgery.
  
=== Unsolved Mathematics & Formal Verification ===
+
Implantable BCIs remain experimental, and significant questions involving safety, reliability, privacy, security, consent, and long-term medical effects must be addressed. An always-on interface that sees what you see and hears what you hear also creates a surveillance surface with no precedent, affecting bystanders who never consented.
After hitting gold-medal performance in the International Mathematical Olympiad, AI is setting its sights on creating original proofs for open mathematical problems. To make sure the math is rock-solid, these systems use formal verification tools like Lean. Imagine a spell-checker, but instead of catching typos, it verifies complex logic step-by-step to guarantee accuracy.
 
  
<youtube>-aXkChrQl1I</youtube>
+
=== 14b. The Larger Moonshot ===
  
=== Reliability & Alignment ===
+
The deeper goal isn't simply to build an AI that is smarter than an individual person.
Continual learning and long-term reliability are the unglamorous but necessary hurdles we still need to clear. Right now, current models tend to lose focus or degrade when they run on their own for too long. If we want AI agents to operate stably over weeks or months, fixing this drift is essential. Think of it like maintaining focus during a marathon rather than just running a quick sprint.
 
  
<youtube>Rs7AdHTc3TI</youtube>
+
It is to create a partnership in which powerful machine intelligence '''extends what an individual can perceive, understand, create, and accomplish'''.
  
= Need to 'Learn' the Wide World Web =
+
If successful, personal superintelligence could change the effective capabilities of an individual:
[https://www.youtube.com/results?search_query=learn+wide+world+web+artificial+intelligence Youtube search...]
 
[https://www.google.com/search?q=learn+wide+world+web+artificial+intelligence ...Google search]
 
  
<youtube>_V8MGzvJEXY</youtube>
+
:'''''One person + powerful AI → capabilities that once required an organization.'''''
<youtube>XZmGGAbHqa0</youtube>
 
  
= Discussions on the Future of AI =
+
Combined with increasingly natural interfaces — from conversation to smart glasses and perhaps eventually brain-computer interfaces — AI could evolve from something people occasionally consult into a persistent cognitive partner.
  
<youtube>XWGnWcmns_M</youtube>
+
That creates one of the major AI moonshots of the coming decade:
  
= <span id="Meeting the Winograd Schema Challenge (WSC)"></span>Meeting the Winograd Schema Challenge (WSC) =
+
:'''''Can superintelligence become a technology that empowers billions of individual people rather than a capability controlled primarily by a small number of institutions?'''''
[https://www.youtube.com/results?search_query=Winograd+Schema+Challenge+WSC+turing+test+machine+artificial+intelligence+ai Youtube search...]
 
[https://www.google.com/search?q=Winograd+Schema+Challenge+WSC+turing+test+machine+artificial+intelligence+ai ...Google search]
 
  
The Winograd Schema Challenge (WSC) is a natural language understanding task proposed as an alternative to the Turing test in 2011. In this work we attempt to solve WSC problems by reasoning with additional knowledge. By using an approach built on top of graph-subgraph isomorphism encoded using Answer Set Programming (ASP) we were able to handle 240 out of 291 WSC problems. The ASP encoding allows us to add additional constraints in an elaboration tolerant manner. In the process we present a graph based representation of WSC problems as well as relevant commonsense knowledge. [https://www.cambridge.org/core/journals/theory-and-practice-of-logic-programming/article/using-answer-set-programming-for-commonsense-reasoning-in-the-winograd-schema-challenge/5CB1C6B33E940A98E086F2EBECA24A09 "Using Answer Set Programming for Commonsense Reasoning in the Winograd Schema Challenge" | Arpit Sharma]
 
 
{|<!-- T -->
 
{|<!-- T -->
 
| valign="top" |
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
{| class="wikitable" style="width: 550px;"
||
+
|| <youtube>EAZe8wDnAOA</youtube> <b>Meta Connect 2025 Keynote</b><br>
<youtube>m3vIEKWrP9Q</youtube>
+
Mark Zuckerberg presents Meta's direction for AI, smart glasses, virtual and augmented reality, and increasingly personal computing. The presentation illustrates Meta's strategy of combining powerful AI with devices that can accompany people throughout their daily lives.
<b>The Sentences Computers Can't Understand, But Humans Can
 
</b><br>The Winograd schema is a language test for intelligent computers. So far, they're not doing well.  
 
 
|}
 
|}
 
|<!-- M -->
 
|<!-- M -->
 
| valign="top" |
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
{| class="wikitable" style="width: 550px;"
||
+
|| <youtube>FASMejN_5gs</youtube> <b>Neuralink Update, Summer 2025</b><br>
<youtube>nZxHF-ndvFo</youtube>
+
Neuralink presents progress in implantable brain-computer interfaces, including neural decoding, computer control, robotic systems, and its longer-term roadmap. It provides a useful view of how direct neural interfaces could eventually complement increasingly capable personal AI.
<b>The Winograd Schema Challenge - Models of Reasoning
 
</b><br>This video corresponds to the online presentation assingment of the subject Models of Reasoning. Authors: Carla Fernández González and Teresa Grau Mateo.  Taking the name from Terry Winograd, who first presented an example following the schema [13], Levesque, Davis and Morgenstern created Winograd Schemas as an alternative to the Turing Test and started a competition to encourage researchers to work in this area of commonsense reasoning.
 
|}
 
|}<!-- B -->
 
{|<!-- T -->
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
||
 
<youtube>GHyDAs2xOos</youtube>
 
<b>Improving Winograd Schemas Using Ambiguous [[context]]s
 
</b><br>I created this presentation for a graduate class at NYU. It also serves as a gentle introduction to Winograd Schemas.  There's one error: at 7:35 I say "she's the receiver of the thanks" when I mean "receiver of the help."  Also, I mean no ill will to Tom Scott. He's a great computer enthusiast and content creator.
 
|}
 
|<!-- M -->
 
| valign="top" |
 
{| class="wikitable" style="width: 550px;"
 
||
 
<youtube>Ey_trzJPp_I</youtube>
 
<b>ICLP19 paper "Using ASP for Commonsense Reasoning in the Winograd Schema Challenge"
 
</b><br>This video is a presentation which provides an overview of the ICLP 2019 conference paper titled [https://www.cambridge.org/core/journals/theory-and-practice-of-logic-programming/article/using-answer-set-programming-for-commonsense-reasoning-in-the-winograd-schema-challenge/5CB1C6B33E940A98E086F2EBECA24A09 "Using Answer Set Programming for Commonsense Reasoning in the Winograd Schema Challenge"]
 
 
|}
 
|}
 
|}<!-- B -->
 
|}<!-- B -->
 +
 +
== 15. Space & Astronomy ==
 +
 +
[https://www.youtube.com/results?search_query=AI+space+exploration+astronomy+autonomous+spacecraft+telescope+data YouTube search...]
 +
[https://www.google.com/search?q=AI+space+exploration+astronomy+autonomous+spacecraft ...Google search]
 +
 +
A page named after the Apollo program should say something about space.
 +
 +
Space is an unusually good fit for autonomous AI for a simple physical reason: '''light-lag makes teleoperation impossible at distance'''. A rover on Mars cannot be joysticked in real time; a probe in the outer solar system even less so. Autonomy is not a convenience there, it is a requirement.
 +
 +
:'''''Commanded from Earth → semi-autonomous → decides what is interesting → decides what to investigate next'''''
 +
 +
Moonshot targets include:
 +
 +
* '''Autonomous science selection''' — spacecraft deciding which observations are worth the bandwidth to transmit.
 +
* '''Survey-scale discovery''' — finding transients, exoplanets, and anomalies in data volumes far beyond human inspection.
 +
* '''Autonomous landing, navigation, and fault recovery''' without ground intervention.
 +
* '''In-space manufacturing and assembly''' by robotic systems.
 +
* '''Mission design''' — AI-optimized trajectories and architectures.
 +
* '''SETI and anomaly detection''' — knowing what an unfamiliar signal looks like.
 +
 +
Space is also where embodied intelligence, world models, and long-horizon reliability all have to work simultaneously, with no possibility of a human taking over.
 +
 +
== 16. Decoding Non-Human Communication ==
 +
 +
[https://www.youtube.com/results?search_query=AI+animal+communication+whale+language+Project+CETI+Earth+Species YouTube search...]
 +
[https://www.google.com/search?q=AI+decoding+animal+communication+whale+language ...Google search]
 +
 +
Machine learning is being applied to the vocalizations of whales, elephants, primates, and birds, searching for structure that human listeners cannot detect — repeated units, combinatorial patterns, context-dependent meaning.
 +
 +
:'''''Record → detect units → find structure → map to context → attempt exchange'''''
 +
 +
This is a genuine moonshot in the Apollo sense: high-risk, uncertain payoff, and transformative if it works. It would be the first time humans communicated with a non-human intelligence, and it would tell us something about how much of language is specifically human.
 +
 +
It is also methodologically humbling. The core difficulty is not pattern detection but '''grounding''' — knowing that a detected pattern means anything at all, rather than being a statistical artifact of an analysis that will find structure in noise if pushed hard enough. Any claim here deserves the same scrutiny applied to AI-generated science in Moonshot 3b.
 +
 +
== 17. Understanding Intelligence Itself ==
 +
 +
[https://www.youtube.com/results?search_query=AI+neuroscience+connectomics+brain+decoding+whole+brain+simulation YouTube search...]
 +
[https://www.google.com/search?q=AI+neuroscience+connectomics+neural+decoding+brain+simulation ...Google search]
 +
 +
[[#14a. From Screens to Continuous AI Interfaces|Moonshot 14a]] covers the brain as an ''interface''. This is the brain as a ''subject''.
 +
 +
We built systems that reason without understanding how the original works. Closing that gap runs in both directions: neuroscience informs AI architectures, and AI makes previously intractable neuroscience possible.
 +
 +
:'''''Map the brain → model the brain → decode the brain → explain intelligence'''''
 +
 +
Targets include:
 +
 +
* '''Connectomics''' — reconstructing complete neural wiring diagrams from electron microscopy, a task made feasible only by machine segmentation.
 +
* '''Neural decoding''' — reconstructing perceived images, intended speech, or inner language from brain activity.
 +
* '''Whole-organism simulation''' — starting with small nervous systems and scaling.
 +
* '''Theories of learning''' that explain both biological and artificial systems.
 +
* '''Mechanisms of memory, attention, and consolidation''' that might transfer to Moonshot 7.
 +
 +
Decoding is where this becomes ethically urgent. Systems that reconstruct mental content from neural signals raise questions of '''cognitive liberty''' — mental privacy, and the right not to have inner states read — that current law does not address.
 +
 +
= Discussions on the Future of AI =
 +
 +
The following discussions provide broader perspectives from leaders working on different versions of the AI moonshot.
 +
 +
<youtube>XWGnWcmns_M</youtube>
 +
 +
<youtube>ZpUKNYcgM-E</youtube>
 +
 +
<youtube>qRQ06azN5FU</youtube>
 +
```
 +
 +
Three notes on what changed beyond additions: I corrected the Sakana Nature DOI to `s41586-026-10265-5`, replaced the non-resolving congress.gov link with the Senate one-pager PDF, and added a disambiguation line for Moonshot AI the company inside Moonshot 23. The AI-as-scientist and automated-researcher sections now carry the self-grading caveat rather than reporting the claims flat.

Latest revision as of 08:41, 10 September 2026

YouTube search... ... Quora search ...Google search ...Google News ...Bing News

The “Sputnik” moment for China came a year ago when a Google computer program, AlphaGo, beat the world’s top master of the ancient board game of Go. Now, China is racing to become the world leader in artificial-intelligence. The term "moonshot" is derived from the Apollo program, which was a series of space missions undertaken by the United States in the 1960s and early 1970s with the goal of landing humans on the Moon. The Apollo program was considered a moonshot because it represented a major technological and engineering challenge that required significant innovation and investment. In context, what do you think would be a "Moonshot" response?



The "moonshot" milestones along the road to Artificial General Intelligence (AGI)


In the context of AI, a "moonshot" refers to a project or goal that aims to achieve a major breakthrough in artificial intelligence that has the potential to transform society or address significant global challenges. Moonshots are high-risk, high-impact goals where advanced AI is used to tackle problems that previously seemed decades away—such as autonomous scientific discovery, general-purpose robotics, radical health extension, new energy technologies, and eventually AI systems that improve AI itself. In short, a moonshot is a high-risk, high-impact goal that aims to achieve a major breakthrough in machine intelligence or use advanced AI to address problems that previously appeared decades away.

AI moonshots fall into two overlapping categories:

  • Capability Moonshots — creating AI systems that can understand the world, reason, learn continuously, conduct research, operate autonomously, improve other AI systems, and act effectively in the physical world.
  • Civilization-Scale Moonshots — applying those capabilities to scientific discovery, medicine, longevity, energy, materials, climate, education, and other major human challenges.

The frontier therefore extends beyond simply making a more capable chatbot. The deeper question is:

What becomes possible when machine intelligence can understand, discover, invent, learn, act, and improve at increasingly general levels?

And immediately behind it, a second question that the first cannot answer on its own:

What has to be true — physically, economically, and institutionally — for that to happen safely?

Calibration: Moonshots Already Achieved

Before surveying what remains, it is worth recording what has already fallen. Each of these was, in its time, described as a definitive test of machine intelligence. Each was reinterpreted as "not real intelligence" shortly after it was accomplished — a pattern sometimes called the AI effect.

Milestone Roughly When How It Was Reinterpreted
Checkers solved 1990s–2007 Search, not thought
Chess grandmaster defeat 1997 Brute force, not understanding
Jeopardy! championship 2011 Retrieval, not comprehension
Go world champion defeat 2016 Narrow domain with clean rules
Protein structure prediction 2020–2024 Pattern-matching on a solved-enough problem
Competition-level mathematics (IMO gold) 2025 Exercises with known answers
Three-party Turing test 2025–2026 Imitation and substitutability, not intelligence

The Turing test case is instructive. In a peer-reviewed PNAS study, GPT-4.5 given a persona prompt was judged human 73% of the time — more often than the actual humans it was paired against — while a 1960s ELIZA control scored 23%, confirming that interrogators were not simply guessing. Yet the same models did not robustly pass without the persona prompt, and in some conditions performed no better than ELIZA. The authors frame the test as a measure of substitutability rather than intelligence, and warn that the near-term consequence is "counterfeit people" — erosion of online trust and displacement of genuine human contact.

Two lessons for everything below:

  • Benchmarks fall faster than expected, and mean less than advertised. Plan for both.
  • Prompting, scaffolding, and tooling often carry more of a result than the underlying model. Always ask what was wrapped around the system before the milestone was declared.

The Moonshot Index

# Moonshot Category Core Question
1 World Models Capability Can AI simulate futures before acting?
2 Recursive Self-Improvement Capability Can AI improve AI?
3 Autonomous Research & Discovery Capability Can AI run the whole discovery cycle?
4 Autonomous Invention & Engineering Capability Can AI build what we haven't?
5 Unsolved Mathematics & Formal Verification Capability Can AI discover new mathematics?
6 Embodied General Intelligence Capability Can AI act in the messy physical world?
7 Lifelong Learning Capability Can AI keep learning after training?
8 Multi-Agent Ecosystems Capability What happens when agents deal with agents?
9 Radical Efficiency Capability Can intelligence run on 20 watts?
10 Biology, Medicine & Longevity Civilization Can AI repair and redesign living systems?
11 Materials & Energy Civilization Can AI find technologies we never found?
12 Weather, Climate & Earth Systems Civilization Can AI forecast and manage a planet?
13 Education Civilization A world-class tutor for every learner?
14 Personal Superintelligence Civilization Superintelligence for individuals, not just institutions?
15 Space & Astronomy Civilization Can AI explore where we cannot go?
16 Decoding Non-Human Communication Civilization Can AI let us talk to other species?
17 Understanding Intelligence Itself Civilization Can AI explain the brain that built it?
18 Compute, Energy & Hardware Substrate Enabling Can we build the launch pad?
19 Evaluation & Measurement Enabling How would we know we succeeded?
20 Interpretability Guardrail Can we see inside the systems we build?
21 Reliability, Alignment & Control Guardrail Can powerful AI stay correctable?
22 Misuse, Dual-Use & Security Guardrail Can capability be released without arming the wrong hands?
23 Governance, Geopolitics & Access Guardrail Who decides, and who benefits?
24 Economics & Labor Guardrail What happens to work?
25 Moral Status of AI Open Does any of this matter to the systems themselves?

Part I — Capability Moonshots

1. World Models — AI That Can Simulate Possible Futures

YouTube search... ...Google search

Humans don't merely react to the world. We continuously construct internal models of it — anticipating what might happen next, imagining alternatives, and considering the consequences of our actions before we act.

World models attempt to give AI a similar capability.

Rather than only recognizing an image, generating text, or responding to the present situation, a world model can represent how an environment changes through time and predict how different actions could alter what happens next.

The progression is:

Perception → World Model → Imagine Futures → Choose Action

Google DeepMind describes its Genie 3 system as a general-purpose world model capable of generating interactive environments in real time. These environments can be used by AI agents to learn how situations evolve and how their own actions affect the environment. DeepMind considers world models an important step toward AGI. "Genie 3 — A New Frontier for World Models" | Google DeepMind

World models could eventually allow AI systems to rehearse actions internally before attempting them in the real world — similar to imagining several possible moves before making a decision.

This capability is especially important for Agents, Robotics, autonomous vehicles, planning, scientific simulation, and long-horizon decision making.

It also makes possible a deeper form of reasoning:

What will happen? → What could happen? → What happens if I intervene? → Which future should I try to create?

Open problems: consistency over long horizons, physical accuracy versus visual plausibility, and whether a model that generates convincing video has actually learned physics or merely learned what physics looks like.

2. Recursive Self-Improvement

YouTube search... ...Google search

The topic of recursive self-improvement is a significant threshold in AI development. Currently, much of the process is still driven by human software engineers and researchers who generate training data, run ablations to test data quality, design experiments, and evaluate models against benchmarks.

However, many laboratories are working to close more of this loop through automation:

  • Automated Judges: Different models can evaluate the quality and correctness of outputs.
  • Generative Feedback: Models can generate new, high-quality training examples and feedback.
  • Adversarial Reasoning: Advanced models can compare, criticize, test, and filter candidate solutions.
  • Automated Experimentation: Agents can modify code, run experiments, evaluate results, and propose subsequent experiments.

The resulting cycle increasingly resembles:

Generate → Evaluate → Select → Train → Test → Improve → Repeat

By feeding successful results back into AI research and post-training, development speed could increase substantially.

More advanced forms of this process could eventually allow AI systems to contribute directly to the design of their successors.

AI helps improve AI → improved AI becomes better at improving AI → the research cycle accelerates.

This possibility is sometimes associated with an intelligence explosion. Mustafa Suleyman and others caution that such a process would still face substantial limits involving compute, experimentation, infrastructure, reliability, and human control.

How much acceleration? Experts disagree sharply. In a forecasting study summarized in the International AI Safety Report 2026, AI-specialist forecasters gave a median 20% probability that the next few years could compress six years of prior advancement into two, while generalist superforecasters put it at 8%. Estimates rose for scenarios in which AI systems outperform humans on month-long research projects. Empirical evidence on whether research automation is actually accelerating progress remains mixed. See also Moonshot 19, since this claim is unusually hard to measure from outside a frontier lab.

3. Autonomous AI Research & Discovery

The frontier of AI is shifting from static chat interfaces toward increasingly autonomous agentic workflows. Rather than performing only one prompt-response cycle, agents can plan tasks, use tools, write and execute code, evaluate results, revise their approach, and coordinate with other agents.

Scientific research is becoming one of the most important proving grounds for these capabilities.

The larger moonshot is:

Can AI participate in the entire discovery cycle — from asking the question to proposing the hypothesis, designing the experiment, interpreting the evidence, and deciding what to try next?

3a. The Automated Researcher

By September 2026, OpenAI reported hitting a self-set target: an automated research intern, defined by the company as a system that can carry out well-defined research tasks under human direction, including tasks that would take a skilled researcher a few days. The organization reported logging roughly 3.1 agent-workdays of effort per eight hours of human labor, with a growing share of researchers running four or more agents concurrently. The stated next target is a fully automated AI researcher by March 2028.

Read the definition carefully. Three qualifiers carry the weight: the tasks are well-defined, a human is steering, and the horizon is days rather than months. OpenAI set the goal, wrote the definition, took the measurements, and published the grade; no external party can independently verify it. The company itself notes that hard-to-automate tasks and compute availability could constrain further progress. Treat it as a credible internal measurement of a self-authored target, not as an externally validated milestone.

3b. AI as Scientist

We have moved past seeing AI only as a tool for retrieving scientific information. Increasingly capable systems can help generate hypotheses, review scientific literature, design experiments, analyze results, and critique possible explanations.

Google DeepMind's Co-Scientist, for example, uses multiple Gemini-based agents to generate, debate, criticize, and refine scientific hypotheses. "Co-Scientist: A Multi-Agent AI Partner to Accelerate Research" | Google DeepMind

Sakana AI's "AI Scientist" explores another version of the concept: systems capable of generating research ideas, running computational experiments, analyzing results, and producing scientific papers. The work was published in Nature, and an unedited, fully AI-generated manuscript passed blind human peer review at an ICLR workshop, scoring above the average human acceptance threshold.

The transition is:

AI answers scientific questions → AI helps investigate scientific questions → AI proposes new scientific questions.

The honest caveat: most such systems are still evaluated on paper generation and benchmark optimization rather than on discoveries that change a field. Verification — knowing whether an AI-generated result is actually correct — is emerging as the binding constraint, and has spawned its own benchmark literature.

4. Autonomous Invention & Engineering

Scientific discovery asks what is true?

Engineering asks what can we build?

A further AI moonshot is therefore autonomous invention — systems capable of designing algorithms, software, chips, machines, materials, and other technologies that human engineers have not previously created.

Google DeepMind's AlphaEvolve combines large language models with automated evaluators and an evolutionary search process to discover and optimize algorithms. Its discoveries have been applied to computing infrastructure, chip design, AI training, mathematics, quantum computing, and other technical problems. "AlphaEvolve: How Our Gemini-Powered Coding Agent Is Scaling Impact Across Fields" | Google DeepMind

Similarly, Gemini Deep Think is being explored for real-world engineering problems including semiconductor fabrication and mechanical design. "Gemini Deep Think" | Google DeepMind

The progression is:

Solve our problems → Discover new principles → Invent new technologies

Eventually, AI-assisted engineering could create technologies whose designs are too complex for any individual human to derive unaided. Note the loop this closes with Moonshot 18: AI that designs better chips is AI that improves its own substrate.

5. Unsolved Mathematics & Formal Verification

After reaching gold-medal-level performance on International Mathematical Olympiad problems, frontier AI systems are increasingly being tested on original mathematical research.

The deeper moonshot isn't simply solving difficult exercises whose answers are already known.

It is:

Can AI discover genuinely new mathematics?

Systems can increasingly combine neural reasoning with formal verification tools such as Lean, allowing a proposed proof to be checked step-by-step against rigorous logical rules.

Formal verification acts like a highly sophisticated proof checker: instead of merely judging whether an argument sounds plausible, the system verifies that each logical step follows correctly.

This combination of creative reasoning and machine-verifiable proof could make mathematics one of the first domains in which highly autonomous AI research becomes possible — precisely because the verification problem that limits Moonshot 3b is, uniquely here, solved.

Concrete targets include long-standing open problems, large-scale formalization of existing mathematical literature, and AI-discovered proofs that human mathematicians accept as novel rather than rediscovered.

6. Embodied General Intelligence

YouTube search... ...Google search

Many researchers describe this as a kind of physical Turing test: a robot that can walk into an unfamiliar home and do the laundry.

Embodied General Intelligence is about moving intelligence from the digital world into messy, unpredictable physical environments.

Traditional industrial robots repeat carefully programmed movements in controlled environments. A generally capable robot must instead perceive unfamiliar surroundings, understand objects, predict physical consequences, recover from errors, and determine how to accomplish goals it has never encountered before.

A household robot, for example, can't simply memorize one sequence for folding laundry or loading a dishwasher. Every home, object, obstruction, person, and situation will be different.

The progression is:

See the World → Understand the World → Act in the World → Adapt to the Unexpected

World models are closely connected to this challenge. Robots may eventually learn many physical skills inside simulated environments before transferring those skills to real machines.

6a. The Data Bottleneck

Language models were trained on an internet-scale corpus that already existed. No equivalent corpus exists for physical manipulation. This is the central reason robotics lags language, and it defines the sub-moonshots:

  • Teleoperation at scale — humans puppeting robots to generate demonstration data, which is slow and expensive.
  • Sim2real transfer — training in simulation where data is free, then closing the reality gap.
  • Learning from human video — extracting manipulation knowledge from ordinary footage of people doing things.
  • Robot foundation models — general-purpose policies transferable across different robot bodies rather than one model per machine.
  • Shared fleet learning — every robot's experience improving every other robot.
No data → borrowed data → simulated data → fleet-generated data → self-sustaining physical learning

6b. Autonomous Vehicles

YouTube search... ...Google search

Autonomous vehicles represent an important early example of embodied intelligence operating in the open world.

Driving requires continuous perception, prediction, planning, spatial reasoning, risk assessment, and interaction with unpredictable human behavior.

A successful autonomous vehicle must constantly answer:

What is happening? → What is likely to happen next? → What should I do about it?

Autonomous driving is also the field's most instructive lesson in the gap between demonstration and deployment: impressive prototypes arrived more than a decade before broad commercial operation, and the delay was almost entirely about the last few percent of reliability. Expect the same shape for household robotics and autonomous research.

7. Lifelong Learning — AI That Keeps Learning

YouTube search... ...Google search

Most modern AI systems undergo an enormous training process and are then largely fixed until their developers train or release another version.

Human intelligence works differently.

We continuously accumulate experiences, form memories, acquire skills, revise beliefs, adapt to changing environments, and learn from both success and failure.

A genuinely general intelligence may eventually require a similar ability to learn throughout its operational lifetime.

The progression is:

Static Model → Persistent Memory → Continual Learning → Lifelong Adaptive Intelligence

Such a system could:

  • Remember important experiences over months or years.
  • Learn new skills without requiring complete retraining.
  • Adapt to an individual user or environment.
  • Incorporate new knowledge while preserving older knowledge.
  • Learn from its own successes and failures.
  • Transfer lessons learned in one domain into another.

This sounds straightforward, but it creates difficult technical and safety problems.

New learning must not erase important existing capabilities — a problem associated with catastrophic forgetting. An adaptive system must also avoid gradually drifting away from its intended goals, learning harmful behaviors, incorporating false information, or modifying itself in ways humans no longer understand.

Lifelong learning therefore connects directly to Memory, Agents, recursive self-improvement, personal superintelligence, and AI alignment.

The moonshot is:

An intelligence that doesn't simply arrive trained — it continues growing through experience.

8. Multi-Agent Ecosystems — When Agents Deal With Agents

YouTube search... ...Google search

Most discussion of agents imagines one agent serving one person. The more consequential scenario is many agents interacting with each other, often without a human in the loop for any individual exchange.

This is already emerging: agents negotiating with agents, agents transacting on behalf of users, agents delegating subtasks to specialist agents, and interoperability protocols that let systems from different vendors discover and call one another.

One agent, one task → agent teams → agent markets → machine-to-machine economies

The moonshot is a functioning agent ecosystem with:

  • Identity and authentication — knowing which agent acts for whom.
  • Delegation and authority limits — bounded permission to spend, commit, or act.
  • Payment and settlement rails designed for machine counterparties.
  • Dispute resolution and liability when an agent errs on someone's behalf.
  • Interoperability standards so agents can discover and use each other's tools.

The risks are distinctive and poorly studied: collusion between agents, cascading failures, flash-crash dynamics at machine speed, and emergent behavior no individual designer intended. Multi-agent dynamics are a live gap in the alignment literature, which has focused mostly on single systems.

9. Radical Efficiency — Intelligence on 20 Watts

YouTube search... ...Google search

The human brain performs general reasoning on roughly 20 watts. Frontier AI systems require datacenters. That gap is not a footnote — it is one of the clearest measures of how far current approaches are from the thing they are imitating.

Efficiency is often treated as an engineering detail rather than a moonshot. It is a moonshot, and possibly the one with the shortest path to broad human impact, because it determines:

  • Whether capable AI runs on a phone, a hearing aid, or a robot rather than in a distant datacenter.
  • Whether personal superintelligence is affordable for billions of people or a subscription for the wealthy.
  • Whether inference cost stops constraining how much reasoning a system is allowed to do.
  • Whether AI's energy footprint is politically and environmentally sustainable.
Frontier capability in a datacenter → in a rack → on a laptop → on a phone → on 20 watts

The competitive dimension is real. Chinese laboratories have pushed hard on the cost-per-capability frontier: Moonshot AI's Kimi K3, at roughly 2.8 trillion parameters, was released as an open-weight model claimed to match or exceed leading Western systems on mainstream benchmarks at a fraction of the running cost — a release that moved markets and prompted a reassessment of assumed US leads. Efficiency, not raw scale, is increasingly where the race is run.

Part II — Civilization-Scale Moonshots

10. AI for Biology, Medicine & Longevity — Understand and Engineer Life

YouTube search... ...Google search

Biology is extraordinarily complex. A living organism contains interacting systems spanning molecules, proteins, genes, cells, organs, environments, and behavior.

AI offers the possibility of modeling these systems at scales that would be extremely difficult for humans to analyze unaided.

AlphaFold demonstrated the potential by transforming protein-structure prediction. The larger moonshot is to move progressively from describing biology toward predicting and eventually designing biological outcomes.

The progression could be:

Understand Biology → Predict Biology → Design Biology → Prevent or Reverse Disease

OpenAI's GPT-Rosalind is a purpose-built frontier reasoning model for life-sciences research, covering molecules, proteins, genes, pathways, and disease biology, with emphasis on multi-step tool use across literature review, sequence-to-function interpretation, experimental planning, and data analysis. It is deployed as a research preview through a trusted-access program with biopharma and research partners, and is evaluated against an internally designed, externally expert-judged benchmark spanning six workflow areas of life-sciences research. "Introducing GPT-Rosalind for Life Sciences Research" | OpenAI

Google DeepMind's AI Co-Scientist similarly explores the use of multi-agent reasoning to generate and refine biological hypotheses. "Co-Scientist" | Google DeepMind

Possible long-term moonshots include:

  • Designing new medicines rather than screening primarily from known compounds.
  • Predicting how genetic changes affect cells and organisms.
  • Modeling complete cells — the virtual cell — and eventually larger biological systems.
  • Creating personalized treatments based on an individual's biology.
  • Detecting disease before symptoms appear.
  • Understanding the biological mechanisms of aging.
  • Extending healthy human lifespan.

The rate limiter is rarely the idea. Target-to-approval for a new drug in the United States takes roughly 10–15 years, and gains at the earliest discovery stages compound downstream. AI that shortens hypothesis generation still runs into wet-lab throughput, clinical trial duration, and regulatory process — which is why closed-loop robotic laboratories (Moonshot 11) matter as much as better models.

The ultimate question is profound:

Can intelligence understand living systems deeply enough to deliberately repair, redesign, and preserve them?

This is also the page's sharpest dual-use case. The same capabilities that accelerate drug discovery lower barriers to engineered biological harm. Frontier laboratories have responded with gated access, biodefense programs, and trusted-deployment models rather than open release. See Moonshot 22.

11. AI for Materials & Energy — Discover Technologies Humans Haven't Found

YouTube search... ...Google search

Human civilization has repeatedly been transformed by the discovery of new materials and new sources of energy — bronze, steel, concrete, semiconductors, petroleum, nuclear energy, photovoltaic materials, lithium-ion batteries, and many others.

Yet discovering useful materials remains an enormous search problem. The number of possible chemical structures is vastly larger than scientists can test experimentally.

AI can search these spaces computationally.

Google DeepMind's GNoME system predicted millions of previously unknown crystal structures, including hundreds of thousands of candidates predicted to be stable. Potential applications include batteries, electronics, superconductors, solar technologies, and many other fields. "Millions of New Materials Discovered with Deep Learning" | Google DeepMind

The discovery process can increasingly connect AI with automated laboratories:

AI predicts → Robots synthesize → Instruments test → Results return to AI → AI proposes the next experiment

This creates another self-improving scientific loop. It is worth noting the standing critique: a predicted stable structure is not a synthesized material, and the gap between computational candidates and validated, manufacturable substances remains large and contested.

Energy represents an equally ambitious target.

Google DeepMind has applied machine learning and reinforcement learning to the extraordinarily difficult problem of controlling plasma inside fusion reactors. In 2025, DeepMind and Commonwealth Fusion Systems announced a partnership using AI simulation, optimization, and real-time control to help advance practical fusion energy. "Bringing AI to the Next Generation of Fusion Energy" | Google DeepMind

The larger chain could be:

AI → New Materials → Better Batteries / Solar / Computing → Fusion → Abundant Clean Energy

The moonshot isn't simply making today's technology more efficient.

It is using AI to discover entirely new technologies that humans haven't yet found — and, circularly, to power the datacenters that the rest of this page depends on.

12. Weather, Climate & Earth Systems

YouTube search... ...Google search

Machine-learning weather models are among the clearest already-realized wins in AI for science — producing forecasts competitive with or better than traditional numerical weather prediction, at a small fraction of the computational cost, in seconds rather than hours on supercomputers.

That success reframes a larger moonshot: modeling the Earth as a coupled system.

Forecast the weather → forecast extremes → model the climate → evaluate interventions

Sub-goals include:

  • Extreme-event prediction with enough lead time to evacuate and prepare.
  • Kilometer-scale regional climate projection usable for actual infrastructure planning.
  • Wildfire, flood, and drought forecasting integrated with response systems.
  • Grid optimization and renewable forecasting — matching intermittent supply to demand.
  • Carbon capture materials discovered through the pipeline in Moonshot 11.
  • Earth digital twins that let policymakers test interventions before committing to them.

This moonshot is unusual in having a genuinely near-term humanitarian payoff, an existing track record, and comparatively low dual-use risk. It is also where AI's own energy footprint becomes uncomfortable to ignore.

13. Education — A World-Class Tutor for Every Learner

YouTube search... ...Google search

Benjamin Bloom's two-sigma problem observed that students receiving one-on-one tutoring performed roughly two standard deviations better than those in conventional classrooms — a result no scalable intervention had matched, because individual human tutors cannot be provided to every learner on Earth.

AI is the first technology with a plausible claim to that scale.

Content delivery → adaptive practice → responsive tutoring → a patient expert who knows this learner

The moonshot requires more than a chatbot that answers homework questions:

  • Modeling the learner — knowing what they actually understand versus what they can recite.
  • Productive struggle — withholding the answer when the struggle is the learning.
  • Motivation and persistence, not just explanation.
  • Assessment integrity in a world where any assignment can be completed by machine.
  • Working in every language, including those with little training data.
  • Teacher augmentation rather than replacement.

The failure mode is specific and already visible: a system that supplies answers efficiently can produce fluent students who have learned less than they appear to have. Measuring genuine learning gains — not engagement, not satisfaction — is the hard part, and connects directly to Moonshot 19.

14. Personal Superintelligence — Powerful AI for Every Individual

YouTube search... ...Google search

Personal superintelligence describes a future in which extremely capable AI isn't concentrated only in governments, corporations, or research laboratories, but is available directly to individuals.

Instead of simply answering questions, a personal AI could understand a person's goals, preferences, history, surroundings, and ongoing activities; reason about complex problems; and take actions on that person's behalf.

Meta has made this idea a central part of its long-term AI strategy. Mark Zuckerberg describes the goal as giving everyone access to a personal superintelligence that can help people create, learn, communicate, pursue their interests, and accomplish things that would previously have required teams of specialists. "Personal Superintelligence" | Meta

The concept represents a shift:

AI as a tool → AI as an assistant → AI as an agent → AI as a personal superintelligence.

A sufficiently capable personal AI might continuously work with an individual across many areas of life — researching questions, teaching new skills, creating software and media, organizing information, communicating with other systems, monitoring projects, and coordinating other specialized AI agents.

Combined with lifelong learning and persistent memory, such a system could increasingly adapt to the individual rather than requiring the individual to continually explain their context to the AI.

This moonshot depends on Moonshot 9. "Personal" superintelligence delivered from a datacenter at frontier inference cost is not personal in any meaningful sense — it is rented, metered, and revocable. Efficiency determines whether this section describes a democratizing technology or a subscription tier.

14a. From Screens to Continuous AI Interfaces

For personal superintelligence to become genuinely personal, the interface between humans and AI may also change.

Today, most people communicate with AI by typing, speaking, or sharing images.

Meta is developing smart glasses as a more continuous interface. Cameras, microphones, displays, and other sensors can allow an AI to see some of what its user sees, hear what the user hears, understand the surrounding context, and provide assistance throughout the day. Meta has suggested that AI-enabled glasses could eventually become a major personal computing platform. "Personal Superintelligence" | Meta

A more radical possibility is the brain-computer interface (BCI).

A BCI establishes a communication pathway between neural activity and a computer. Instead of translating every intention into movements of the hands, eyes, or voice, some information can potentially move directly between the nervous system and a digital device.

Current implantable BCIs are primarily being developed as medical technologies. Neuralink, for example, is testing systems that allow people with severe paralysis to control computers, phones, robotic devices, and other equipment using neural activity. "Two Years of Telepathy" | Neuralink

The longer-term possibility is much broader:

Brain → AI → Digital World

A high-bandwidth BCI combined with increasingly capable personal AI could eventually create a much more direct relationship between human intention and machine intelligence.

The AI might infer what a person wants to accomplish, retrieve information, operate digital systems, or communicate with other AI agents with less dependence on keyboards, screens, or spoken commands.

This does not mean that brain implants are required for personal superintelligence. Wearable devices, voice interfaces, augmented-reality glasses, and other technologies may provide much of this capability without surgery.

Implantable BCIs remain experimental, and significant questions involving safety, reliability, privacy, security, consent, and long-term medical effects must be addressed. An always-on interface that sees what you see and hears what you hear also creates a surveillance surface with no precedent, affecting bystanders who never consented.

14b. The Larger Moonshot

The deeper goal isn't simply to build an AI that is smarter than an individual person.

It is to create a partnership in which powerful machine intelligence extends what an individual can perceive, understand, create, and accomplish.

If successful, personal superintelligence could change the effective capabilities of an individual:

One person + powerful AI → capabilities that once required an organization.

Combined with increasingly natural interfaces — from conversation to smart glasses and perhaps eventually brain-computer interfaces — AI could evolve from something people occasionally consult into a persistent cognitive partner.

That creates one of the major AI moonshots of the coming decade:

Can superintelligence become a technology that empowers billions of individual people rather than a capability controlled primarily by a small number of institutions?
Meta Connect 2025 Keynote

Mark Zuckerberg presents Meta's direction for AI, smart glasses, virtual and augmented reality, and increasingly personal computing. The presentation illustrates Meta's strategy of combining powerful AI with devices that can accompany people throughout their daily lives.

Neuralink Update, Summer 2025

Neuralink presents progress in implantable brain-computer interfaces, including neural decoding, computer control, robotic systems, and its longer-term roadmap. It provides a useful view of how direct neural interfaces could eventually complement increasingly capable personal AI.

15. Space & Astronomy

YouTube search... ...Google search

A page named after the Apollo program should say something about space.

Space is an unusually good fit for autonomous AI for a simple physical reason: light-lag makes teleoperation impossible at distance. A rover on Mars cannot be joysticked in real time; a probe in the outer solar system even less so. Autonomy is not a convenience there, it is a requirement.

Commanded from Earth → semi-autonomous → decides what is interesting → decides what to investigate next

Moonshot targets include:

  • Autonomous science selection — spacecraft deciding which observations are worth the bandwidth to transmit.
  • Survey-scale discovery — finding transients, exoplanets, and anomalies in data volumes far beyond human inspection.
  • Autonomous landing, navigation, and fault recovery without ground intervention.
  • In-space manufacturing and assembly by robotic systems.
  • Mission design — AI-optimized trajectories and architectures.
  • SETI and anomaly detection — knowing what an unfamiliar signal looks like.

Space is also where embodied intelligence, world models, and long-horizon reliability all have to work simultaneously, with no possibility of a human taking over.

16. Decoding Non-Human Communication

YouTube search... ...Google search

Machine learning is being applied to the vocalizations of whales, elephants, primates, and birds, searching for structure that human listeners cannot detect — repeated units, combinatorial patterns, context-dependent meaning.

Record → detect units → find structure → map to context → attempt exchange

This is a genuine moonshot in the Apollo sense: high-risk, uncertain payoff, and transformative if it works. It would be the first time humans communicated with a non-human intelligence, and it would tell us something about how much of language is specifically human.

It is also methodologically humbling. The core difficulty is not pattern detection but grounding — knowing that a detected pattern means anything at all, rather than being a statistical artifact of an analysis that will find structure in noise if pushed hard enough. Any claim here deserves the same scrutiny applied to AI-generated science in Moonshot 3b.

17. Understanding Intelligence Itself

YouTube search... ...Google search

Moonshot 14a covers the brain as an interface. This is the brain as a subject.

We built systems that reason without understanding how the original works. Closing that gap runs in both directions: neuroscience informs AI architectures, and AI makes previously intractable neuroscience possible.

Map the brain → model the brain → decode the brain → explain intelligence

Targets include:

  • Connectomics — reconstructing complete neural wiring diagrams from electron microscopy, a task made feasible only by machine segmentation.
  • Neural decoding — reconstructing perceived images, intended speech, or inner language from brain activity.
  • Whole-organism simulation — starting with small nervous systems and scaling.
  • Theories of learning that explain both biological and artificial systems.
  • Mechanisms of memory, attention, and consolidation that might transfer to Moonshot 7.

Decoding is where this becomes ethically urgent. Systems that reconstruct mental content from neural signals raise questions of cognitive liberty — mental privacy, and the right not to have inner states read — that current law does not address.

Discussions on the Future of AI

The following discussions provide broader perspectives from leaders working on different versions of the AI moonshot.

```

Three notes on what changed beyond additions: I corrected the Sakana Nature DOI to `s41586-026-10265-5`, replaced the non-resolving congress.gov link with the Senate one-pager PDF, and added a disambiguation line for Moonshot AI the company inside Moonshot 23. The AI-as-scientist and automated-researcher sections now carry the self-grading caveat rather than reporting the claims flat.