Difference between revisions of "Moonshots"
(→Meeting the Winograd Schema Challenge (WSC)) |
m |
||
| Line 24: | Line 24: | ||
* [[What is Artificial Intelligence (AI)? | Artificial Intelligence (AI)]] ... [[Generative AI]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]] | * [[What is Artificial Intelligence (AI)? | Artificial Intelligence (AI)]] ... [[Generative AI]] ... [[Machine Learning (ML)]] ... [[Deep Learning]] ... [[Neural Network]] ... [[Reinforcement Learning (RL)|Reinforcement]] ... [[Learning Techniques]] | ||
* [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[Bing/Copilot]] | [[Microsoft]] ... [[Gemini]] | [[Google]] ... [[Claude]] | [[Anthropic]] ... [[Perplexity]] ... [[You]] ... [[phind]] ... [[Ernie]] | [[Baidu]] | * [[Conversational AI]] ... [[ChatGPT]] | [[OpenAI]] ... [[Bing/Copilot]] | [[Microsoft]] ... [[Gemini]] | [[Google]] ... [[Claude]] | [[Anthropic]] ... [[Perplexity]] ... [[You]] ... [[phind]] ... [[Ernie]] | [[Baidu]] | ||
| + | * [[Data Science]] ... [[Data Governance|Governance]] ... [[Data Preprocessing|Preprocessing]] ... [[Feature Exploration/Learning|Exploration]] ... [[Data Interoperability|Interoperability]] ... [[Algorithm Administration#Master Data Management (MDM)|Master Data Management (MDM)]] ... [[Bias and Variances]] ... [[Benchmarks]] ... [[Datasets]] | ||
* [https://www.nature.com/articles/s41586-024-07936-z The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Sakana AI Team - Nature] ...Breakthrough in AI-generated research papers achieving human-level acceptance at workshops. | * [https://www.nature.com/articles/s41586-024-07936-z The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Sakana AI Team - Nature] ...Breakthrough in AI-generated research papers achieving human-level acceptance at workshops. | ||
* [https://www.congress.gov/bill/119th-congress/senate-bill/AI-Grand-Challenges-Act AI Grand Challenges Act | U.S. Senate - 2026] ...Proposed bipartisan legislation to fund prize competitions for AI breakthroughs in health and security. | * [https://www.congress.gov/bill/119th-congress/senate-bill/AI-Grand-Challenges-Act AI Grand Challenges Act | U.S. Senate - 2026] ...Proposed bipartisan legislation to fund prize competitions for AI breakthroughs in health and security. | ||
Revision as of 07:11, 10 September 2026
YouTube search... ... Quora search ...Google search ...Google News ...Bing News
- Artificial General Intelligence (AGI) to Singularity ... Curious Reasoning ... Emergence ... Moonshots ... Explainable AI ... Automated Learning
- Agents ... Robotic Process Automation ... Assistants ... Personal Companions ... Productivity ... Email ... Negotiation ... LangChain
- Large Language Model (LLM) ... Multimodal ... Foundation Models (FM) ... Generative Pre-trained ... Transformer ... Attention ... GAN ... BERT
- In-Context Learning (ICL) ... Context ... Out-of-Distribution (OOD) Generalization
- Immersive Reality ... Metaverse ... Omniverse ... Transhumanism ... Religion
- Telecommunications ... Computer Networks ... 5G ... Satellite Communications ... Quantum Communications ... Communication Agents ... Smart Cities ... Digital Twin ... Internet of Things (IoT)
- Creatives ... History of Artificial Intelligence (AI) ... Neural Network History ... Rewriting Past, Shape our Future ... Archaeology ... Paleontology
- Humor
- Joke-Telling Robots Are the Final Frontier of Artificial Intelligence | Becky Ferreira - Vice ... Humor requires self-awareness, spontaneity, linguistic sophistication, and empathy. Not easy for a robot.
- How to build your own AlphaZero AI using Python and Keras
- Artificial Intelligence (AI) ... Generative AI ... Machine Learning (ML) ... Deep Learning ... Neural Network ... Reinforcement ... Learning Techniques
- Conversational AI ... ChatGPT | OpenAI ... Bing/Copilot | Microsoft ... Gemini | Google ... Claude | Anthropic ... Perplexity ... You ... phind ... Ernie | Baidu
- Data Science ... Governance ... Preprocessing ... Exploration ... Interoperability ... Master Data Management (MDM) ... Bias and Variances ... Benchmarks ... Datasets
- The AI Scientist: Towards Fully Automated Open-Ended Scientific Discovery | Sakana AI Team - Nature ...Breakthrough in AI-generated research papers achieving human-level acceptance at workshops.
- AI Grand Challenges Act | U.S. Senate - 2026 ...Proposed bipartisan legislation to fund prize competitions for AI breakthroughs in health and security.
The “Sputnik” moment for China came a year ago when a Google computer program, AlphaGo, beat the world’s top master of the ancient board game of Go. Now, China is racing to become the world leader in artificial-intelligence. The term "moonshot" is derived from the Apollo program, which was a series of space missions undertaken by the United States in the 1960s and early 1970s with the goal of landing humans on the Moon. The Apollo program was considered a moonshot because it represented a major technological and engineering challenge that required significant innovation and investment. In context, what do you think would be a "Moonshot" response?
The 'moonshot' milestones along the road to Artificial General Intelligence (AGI)
In the context of AI, a "moonshot" refers to a project or goal that aims to achieve a major breakthrough in artificial intelligence that has the potential to transform society or address significant global challenges. Moonshots are high-risk, high-impact goals where advanced AI is used to tackle problems that previously seemed decades away—such as autonomous scientific discovery, general-purpose robotics, radical health extension, new energy technologies, and eventually AI systems that improve AI itself.
Contents
Able to Predict the Future
Youtube search... ...Google search
Can Conjure & Ask Questions
Youtube search... ...Google search
Recursive Self Improvement
The topic of recursive self-improvement is a significant threshold in AI development. Currently, the process is largely driven by human software engineers who manually generate training data, run ablations to test data quality, and evaluate models against benchmarks.
However, many labs are now working to close this loop to automate the process:
- Automated Judges: Different models will act as judges to evaluate the quality of outputs.
- Generative Feedback: Models will generate new, high-quality training data autonomously.
- Adversarial Reasoning: Advanced models will reason over which data to include, effectively filtering for quality.
By feeding this output back into the post-training process, development speed will likely increase significantly. While some speculate this could lead to an intelligence explosion, Mustafa Suleyman notes that achieving this requires substantial compute and, without proper human oversight or control, it introduces significant risks.
The frontier of AI has shifted from static chat interfaces to autonomous "Agentic Workflows." These systems are designed for recursive self-improvement and multi-agent orchestration, moving beyond simple prompt-response cycles.
The Automated Researcher
By September 2026, OpenAI hit a major milestone with its "automated research intern." Think of it as a tireless assistant that handles well-defined research tasks while a human steers the ship. For every eight hours a person puts in, the system completes about three days' worth of work. The ultimate goal is to have a fully autonomous AI researcher running by March 2028.
AI as Scientist
We have moved past seeing AI as just a tool; it is now a true collaborator. Models like Sakana AI's "AI Scientist" are actually producing fresh, peer-review-quality research. They are running experiments in areas like virtual cell modeling and designing new drugs from scratch. It is like having a digital post-doc working around the clock in the lab.
Unsolved Mathematics & Formal Verification
After hitting gold-medal performance in the International Mathematical Olympiad, AI is setting its sights on creating original proofs for open mathematical problems. To make sure the math is rock-solid, these systems use formal verification tools like Lean. Imagine a spell-checker, but instead of catching typos, it verifies complex logic step-by-step to guarantee accuracy.
Reliability & Alignment
Continual learning and long-term reliability are the unglamorous but necessary hurdles we still need to clear. Right now, current models tend to lose focus or degrade when they run on their own for too long. If we want AI agents to operate stably over weeks or months, fixing this drift is essential. Think of it like maintaining focus during a marathon rather than just running a quick sprint.
Embodied General Intelligence
Many people call this the "physical Turing test." Embodied General Intelligence is all about getting robots to navigate messy, unpredictable real-world spaces. Instead of repeating the same motion on an assembly line, these robots need to figure out how to fold laundry or do the dishes in a kitchen they have never seen before.
Autonomous Vehicles
Youtube search... ...Google search
Need to 'Learn' the Wide World Web
Youtube search... ...Google search
Discussions on the Future of AI
Winograd Schema Challenge (WSC): From AI Moonshot to Historical Milestone
YouTube search... ...Google search
The Winograd Schema Challenge (WSC) was proposed by Hector Levesque in 2011 as a more focused alternative to the Turing Test. It tested whether an AI system could resolve ambiguous pronouns in sentences where the correct answer appears to require commonsense reasoning and knowledge about how the world works.
A classic example is:
- The city councilmen refused the demonstrators a permit because they feared violence.
- The city councilmen refused the demonstrators a permit because they advocated violence.
In the first sentence, they normally refers to the city councilmen; in the second, it refers to the demonstrators. Humans generally make this distinction effortlessly by using contextual and commonsense knowledge rather than grammatical rules alone.
When the challenge was introduced, solving problems like these was considered a significant goal for AI. The original schemas were deliberately constructed to resist simple statistical shortcuts, making the WSC an influential test of whether machines could move beyond pattern matching toward genuine language understanding and reasoning. "The Winograd Schema Challenge" | Ernest Davis, Leora Morgenstern, and Charles Ortiz
The status of the challenge has since changed. By 2019, multiple systems based on large pretrained transformer language models and fine-tuned on related tasks were achieving greater than 90 percent accuracy. A later review by Vid Kocijan, Ernest Davis, Thomas Lukasiewicz, Gary Marcus, and Leora Morgenstern concluded that the original WSC had largely been overcome as a benchmark. "The Defeat of the Winograd Schema Challenge" | Artificial Intelligence, 2023
This did not mean that AI had solved commonsense reasoning. Instead, the WSC became an important example of a broader problem in AI evaluation: a system can master a benchmark without necessarily mastering the general capability that the benchmark was intended to measure.
Research published in 2024 reinforced this point. Prompted large language models performed strongly on Winograd-style datasets while performing substantially worse on other forms of pronoun ambiguity. This suggests that performance on any single benchmark provides only a partial picture of a model's underlying reasoning ability. "Solving the Challenge Set without Solving the Task: On Winograd Schemas as a Test of Pronominal Coreference Resolution" | Ian Porada and Jackie C. K. Cheung, 2024
The WSC therefore remains significant today, but primarily as a historical milestone in AI development and evaluation. It helped focus research on commonsense reasoning, contributed to later benchmarks such as SuperGLUE and WinoGrande, and demonstrated an important lesson for modern AI: benchmark success and general intelligence are not the same thing.
Earlier research also illustrates how approaches to the problem evolved. Before large pretrained language models became dominant, researchers explored explicit knowledge representation and symbolic reasoning. One approach used graph representations and Answer Set Programming (ASP) to incorporate additional commonsense knowledge and constraints, handling 240 of 291 WSC problems examined in that work. "Using Answer Set Programming for Commonsense Reasoning in the Winograd Schema Challenge" | Arpit Sharma
|
|