Agents

From
Jump to: navigation, search

YouTube ... Quora ...Google search ...Google News ...Bing News



Agents flip our copilot paradigm: from AI being human's copilot to humans being AI's copilot.

The paradigm of AI copilots is undergoing a transformation, shifting from AI serving as a supportive assistant to humans, to humans taking on the role of AI copilots. This shift represents a significant change in how individuals interact with AI, enabling users to harness the power of AI to enhance decision-making and efficiency. This shift in the copilot paradigm signifies a move towards a more collaborative and empowering relationship between humans and AI, where humans actively engage with AI to augment their capabilities and productivity.


There are 3 stages; Beside, Inside, & Outside - Steven Bathiche

  1. AI will be beside what we are doing: AI is a valuable tool; a helper, a sidebar, an assistant that can aid us in many aspects of our lives. As a conversational AI, it can help us with research, writing, and presentations, breaking our mental blacks, managing our personal lives by scheduling appointments and managing our finances. By automating tasks such as data entry and email processing, AI frees up our time so that we can focus on more important tasks.
  2. AI will inside what we are doing: AI will become even more integrated into our daily lives, understanding our needs and preferences and anticipating our actions. It will be everywhere, acting as the scaffolding for our interactions with technology. Instead of waiting for us to make a keystroke or select an application, AI will take the lead, choosing the best app for us and coaching us through a vast array of possibilities.
  3. AI will be outside what we are doing: AI agents will act as intermediaries, pulling information from disparate data sources and interpreting the world for us. They will help us navigate and understand the context, acting as facilitators and eliminating the challenge of starting with a blank sheet.

An intelligent agent is anything which perceives its environment, takes actions autonomously in order to achieve goals, and may improve its performance with learning or acquiring knowledge. An agent has an "objective function" that encapsulates all the IA's goals. Such an agent is designed to create and execute whatever plan will, upon completion, maximize the expected value of the objective function.[2] For example, a reinforcement learning agent has a "reward function" that allows the programmers to shape the IA's desired behavior,[3] and an evolutionary algorithm's behavior is shaped by a "fitness function". - Wikipedia

What is OpenClaw and related efforts?

OpenClaw is an open-source autonomous AI agent framework capable of executing multi-step tasks across various digital platforms. Created in late 2025 by Austrian developer Peter Steinberger (formerly known as "Clawdbot" and "Moltbot" before a rebranding), the software runs locally on a user's machine and serves as a bridge between Large Language Models (LLMs) and local system tools. Unlike traditional chatbots that only generate text, OpenClaw is designed to perform active workflows—such as managing files, sending emails, and writing code—initiated via messaging apps like WhatsApp, Telegram (software), and Slack (software).

The project gained viral popularity in early 2026, often described as "AI that actually does things," due to its ability to retain persistent memory and execute complex chains of actions without continuous human oversight. However, its deep system access has drawn scrutiny from cybersecurity researchers, who have highlighted potential risks involving unauthorized data modification and "rogue" agent behavior.


Related Efforts

Following OpenClaw's rapid rise, major technology companies accelerated the development of competing autonomous agents or integrated similar capabilities into their ecosystems.

Google

In May 2026, reports emerged that Google was internally testing a competitor codenamed "Remy" (derived from remigus, Latin for "oarsman"). Integrated into a staff-only version of Gemini, Remy is described as a "24/7 personal agent" capable of proactively handling workflows across the Google Workspace ecosystem (Gmail, Calendar, Drive) rather than just responding to prompts. The project is seen as a direct response to OpenClaw's "agentic" capabilities, aiming to offer a more secure, managed alternative for task automation.

OpenAI

Rather than releasing a direct clone, OpenAI "absorbed" the primary talent behind the project. In February 2026, CEO Sam Altman announced the hiring of OpenClaw creator Peter Steinberger to work on OpenAI's next generation of personal agents. While OpenClaw itself remained open-source, this move was widely interpreted as a strategic pivot by OpenAI to internalize the "agentic" workflow infrastructure that Steinberger had popularized.

Microsoft

Microsoft developed an enterprise-focused alternative internally codenamed "Project Lobster" (a nod to the crustacean theme of OpenClaw). This initiative evolved into the official product Microsoft Scout, announced in June 2026. Scout is integrated into Microsoft 365 and Windows, offering similar autonomous capabilities—such as scheduling and data processing—but wrapped in enterprise-grade security and identity management (Entra ID) to address the compliance concerns associated with self-hosted agents.

Anthropic

Early tension existed between OpenClaw and Anthropic due to the original name "Clawdbot" being a homophone for their model "Claude," which led to a legal request for a rename. Anthropic is reportedly developing a remote control feature called "Dispatch" to allow its models to operate agents in the background.

Meta

Meta's entry into the space includes both product development and strategic acquisitions. • Manus: In March 2026, Meta launched a desktop application for its AI agent, Manus (sometimes referred to in leaks as "Manis" or "Malice"), which brings agentic capabilities directly to personal computers. The tool allows users to control desktop applications and local files, positioning it as a consumer-friendly alternative to OpenClaw. • Moltbook Acquisition: Meta acquired Moltbook, a viral social media platform populated exclusively by AI agents (originally built for OpenClaw agents). The acquisition brought the platform's founders into Meta's "Superintelligence Labs" to further develop social behaviors for autonomous agents.

Other Competitors

Salesforce: The company is developing an enterprise alternative named "Albert." Perplexity: In February 2026, Perplexity launched a local agent system dubbed "Personal Computer." Open Source: The OpenClaw ecosystem has spawned numerous forks and lightweight alternatives, including NanoClaw (focused on security sandboxing), Hermes Agent, and ZeroClaw.

What is an AI Agent?

LlamaIndex’s Agent

LlamaIndex's Agent is a large language model (LLM) chatbot developed by LlamaIndex, a company that specializes in alpaca and llama products and information. The agent is trained on a massive dataset of text and code, including LlamaIndex's own database of alpaca and llama information. This allows the agent to answer a wide range of questions about alpacas and llamas, including their biology, behavior, care, and products. The agent can also be used to generate creative text formats of text content, like poems, code, scripts, musical pieces, email, letters, etc. It can also be used to translate languages. LlamaIndex's Agent is a valuable resource for anyone who is interested in learning more about alpacas and llamas. It is also a useful tool for alpaca and llama breeders, farmers, and businesses. Here are some examples of how LlamaIndex's Agent can be used:

  • Answer questions about alpacas and llamas: The agent can answer a wide range of questions about alpacas and llamas, including their biology, behavior, care, and products. For example, you could ask the agent "How long do alpacas live?" or "What is the difference between a llama and an alpaca?"
  • Generate creative text formats: The agent can be used to generate creative text formats of text content, like poems, code, scripts, musical pieces, email, letters, etc. For example, you could ask the agent to write a poem about alpacas or a script for a video about llama care.
  • Translate languages: The agent can also be used to translate languages. For example, you could ask the agent to translate a Spanish article about alpacas into English.

Langchain’s Agent

LangChain's Agent is a type of artificial intelligence (AI) that can be used to automate a variety of tasks. It is a language model that is trained on a massive dataset of text and code. This allows the agent to understand and respond to natural language, and to perform many kinds of tasks, including:

  • Answering questions in a comprehensive and informative way
  • Generating different creative text formats of text content
  • Translating languages
  • Automating tasks that would otherwise require human intervention

LangChain's Agent is still under development, but it has the potential to revolutionize the way we work and live. For example, it could be used to develop new types of customer service chatbots, to automate tasks in software development, and to create new forms of creative content. Here are some examples of how LangChain's Agent could be used in the real world:

  • A customer service chatbot could use LangChain's Agent to answer customer questions in a more comprehensive and informative way than traditional chatbots.
  • A software development team could use LangChain's Agent to automate tasks such as writing unit tests and generating documentation.
  • A writer could use LangChain's Agent to generate ideas for new articles or stories, or to translate their work into other languages.

LangChain's Agent is a powerful tool that has the potential to be used in a wide variety of applications. As the technology continues to develop, we can expect to see even more innovative and exciting uses for LangChain's Agent in the future. In addition to the examples above, LangChain's Agent could also be used to:

  • Develop new educational tools that can personalize instruction and provide feedback to students.
  • Create new types of creative content, such as interactive stories and games.
  • Develop new tools for scientific research, such as helping scientists to design experiments and analyze data.
  • Automate tasks in a variety of industries, such as healthcare, finance, and manufacturing.

DeepMind's WebAgent

YouTube ... Quora ...Google search ...Google News ...Bing News

Google DeepMind's WebAgent can autonomously navigate and interact with websites, following natural language instructions to complete tasks; capable of understanding and interacting with the vast and complex world of the internet.



'https://miro.medium.com/v2/resize:fit:828/format:webp/1*sA904oL8A4DUQ6FQInNAbw.png'


Key Features of WebAgent:

  • Planning and Long Context Understanding: WebAgent can break down complex instructions into smaller, manageable steps and analyze lengthy HTML documents to extract relevant information. This allows it to effectively plan its actions and adapt to the dynamic nature of websites.
  • Large Language Model (LLM): WebAgent's design incorporates two distinct large language models, HTML-T5 and Flan-U-PaLM, each specialized for specific tasks. This modular approach allows for efficient processing and enhanced performance. HTML-T5: A domain-specific LLM trained on a massive corpus of HTML documents, specializing in understanding and summarizing website content. Flan-U-PaLM: A powerful LLM capable of generating executable Python code, enabling WebAgent to interact with websites directly.
  • Program Synthesis: WebAgent employs program synthesis techniques to translate natural language instructions into executable code, allowing it to perform actions on websites. WebAgent can generate executable Python code based on its understanding of instructions and website structure. This enables it to directly interact with websites, filling forms, clicking buttons, and extracting data.
  • Reinforcement Learning (RL): Reinforcement learning algorithms are used to train WebAgent, enabling it to learn from experience and improve its performance over time.
  • Web Scraping Techniques: WebAgent utilizes web scraping techniques to extract relevant information from websites, such as text, images, and links.
  • Natural Language Processing (NLP): NLP techniques are employed to analyze and understand natural language instructions, enabling WebAgent to follow complex commands.
  • Machine Learning (ML): Machine learning algorithms are used to train and optimize WebAgent's various components, enhancing its ability to perform tasks accurately and efficiently.

Potential Applications:

  • Automated Web Testing: WebAgent can be employed to automate website testing, identifying bugs, and ensuring functionality across different platforms and browsers.
  • Web Scraping and Data Extraction: WebAgent's ability to understand and navigate websites makes it an ideal tool for extracting specific data from various sources, streamlining data collection processes.
  • Personalized Web Assistance: WebAgent could serve as a personalized web assistant, helping users with tasks like booking appointments, making online purchases, or managing personal accounts.

Jarvis

YouTube ... Quora ...Google search ...Google News ...Bing News

Jarvis is a project from Microsoft that uses ChatGPT as the controller for a system where it can employ a variety of other models as needed to respond to your prompt. Microsoft Jarvis connects LLMs with ML community. Language serves as an interface for Large Language Model (LLM)s to connect numerous AI models for solving complicated AI tasks. Solving complicated AI tasks with different domains and modalities is a key step toward advanced artificial intelligence. While there are abundant AI models available for different domains and modalities, they cannot handle complicated AI tasks. Considering Large Language Model (LLM) have exhibited exceptional ability in language understanding, generation, interaction, and reasoning, we advocate that LLMs could act as a controller to manage existing AI models to solve complicated AI tasks and language could be a generic interface to empower this.



HuggingGPT is one instance of Jarvis, a web-based chatbot at Hugging Face, an online AI community which hosts thousands of open-source models.



When a user makes a request to the bot, Jarvis plans the task, chooses which models it needs, has those models perform the task and then generates and issues a response. The workflow of this system consists of four stages:

  1. Task Planning: Using ChatGPT to analyze the requests of users to understand their intention, and disassemble them into possible solvable tasks via prompts.
  2. Model Selection: To solve the planned tasks, ChatGPT selects expert models that are hosted on Hugging Face based on model descriptions.
  3. Task Execution: Invoke and execute each selected model, and return the results to ChatGPT.
  4. Response Generation: Finally, using ChatGPT to integrate the prediction of all models, and generate answers for users.


HuggingGPT

YouTube ... Quora ...Google search ...Google News ...Bing News

HuggingGPT that the Microsoft researchers have set up that leverages Large Language Model (LLM) such as ChatGPT to connect various AI models in machine learning communities to solve AI tasks. Specifically, HuggingGPT uses ChatGPT to conduct task planning when receiving a user request, select models according to their function descriptions available in Hugging Face, execute each subtask with the selected AI model, and summarize the response according to the execution results. To use HuggingGPT, you'll need to obtain an OpenAPI API Key if you don't already have one and sign up for a free account at Hugging Face. Once you've logged in to the site, navigate to Settings -> Access Tokens by clicking the links in the left rail.

Auto-GPT

YouTube ... Quora ...Google search ...Google News ...Bing News



Auto-GPT (AutoGPT) is an open-source Python AI Agent application based on GPT-4 that can self-prompt. This means that if the user states an end goal, the system can work out the steps needed to get there and carry them out. Auto-GPT works by setting a goal; the AI will then generate and complete tasks. Basically, it does all the follow-up work for you, asking and answering its own prompts in a loop. It utilizes the GPT-4 API and can perform a task with little human intervention.

Auto-GPT manages short-term and long-term memory by writing to and reading from databases and files; manages context window length requirements with summarization; can perform internet-based actions such as web searching, web form, and API interactions unattended; and includes text-to-speech for voice output. However, it has limitations in understanding and retaining extensive contextual information because the GPT model it leverages has a token limit. One way to address this contextual issue is to access a window of historical messages, such as the last ten messages or a fixed number of tokens, without exceeding the token limit of a single conversation. However, this method restricts Auto-GPT from accessing earlier contextual information, which might lead to the failure of Auto-GPT to accomplish its goal

To install Auto-GPT, you will need to have Python and Pip installed on your computer. You can download the latest version of Python from the official website and install it on your computer. You will also need to add API keys to use Auto-GPT. You can go to the GitHub release page of Auto-GPT and download the ZIP file by clicking on “Source code (zip)”

If you want Auto-GPT to speak using ElevenLabs, you will need to have an ElevenLabs API key. You can obtain your ElevenLabs API key from their website. Once you have your API key, you can add it to the .env file in the Auto-GPT directory.



Auto-GPT is not yet capable of achieving the AGI (Artificial General Intelligence) due to data quality, generalization, and explainability issues. - Kanwal Mehreen



Breaking Down AutoGPT | Kanwal Mehreen - KDnuggets ... here are the steps

  1. Input from the User
  2. Task Creation Agent
  3. Task Prioritization Agent
  4. Communication Between Agents
  5. Final Result - It also uses external memory to keep track of history and learn from its past experiences to generate more precise results.

The actions of these agents are visible on the user end in the following form:

  • Thoughts: AI agent share their thoughts after completing the action
  • Reasoning: It explains its choices of why is it choosing a particular course of action
  • Plan: The plan includes the new set of tasks
  • Criticism: Critically review the choices by identifying the limitations or concerns



Reflexion

YouTube ... Quora ...Google search ...Google News ...Bing News

Reflexion is a meta-technique approach that endows an agent with dynamic memory and self-reflection capabilities to enhance its existing reasoning trace and task-specific action choice abilities1 It builds on recent research and allows agents to learn from their mistakes and solve novel problems efficiently through a process of trial and error. Reflexion’s success has been demonstrated through evaluations in AlfWorld and HotPotQA environments, achieving success rates of 97% and 51%, respectively. To achieve full automation, Reflexion introduces a straightforward yet effective heuristic that enables the agent to pinpoint hallucination instances, avoid repetition in action sequences, and, in some environments, construct an internal memory map of the given environment.


Self-reflection allows humans to efficiently solve novel problems through a process of trial and error.




Autonomous GPT

YouTube ... Quora ...Google search ...Google News ...Bing News



AgentGPT

YouTube ... Quora ...Google search ...Google News ...Bing News

AgentGPT allows you to configure and deploy Autonomous AI agents. Name your custom AI and have it embark on any goal imaginable. It will attempt to reach the goal by thinking of tasks to do, executing them, and learning from the results

BabyAGI

YouTube ... Quora ...Google search ...Google News ...Bing News


OpenAGI

OpenAGI is an open-source research platform for artificial general intelligence (AGI). It's designed to offer complex, multi-step tasks. OpenAGI offers complex tasks, datasets, metrics, and models for solving tasks. It also provides task-specific datasets, evaluation metrics, and a diverse range of extensible models. OpenAGI uses a dual strategy, integrating standard benchmark tasks for benchmarking and evaluation, and open-ended tasks including more expandable models, tools, plugins, or APIs for creative problem-solving. OpenAGI formulates complex tasks as natural language queries, serving as input to the LLM. It also proposes an LLM+RLTF approach to learning better task design.


AutoGen

YouTube ... Quora ...Google search ...Google News ...Bing News

AutoGen Studio 2.0** is a comprehensive tool suitable for developers of all skill levels. It simplifies AI development by providing an intuitive interface and extensive toolset. The platform caters to a broad spectrum of developers, from novices to experts.

Getting Started:

  • Environment Preparation: Crucial steps include installing **Python** and **Anaconda**.
  • Configuring LLM Provider: Obtain an API key from **OpenAI** or **Azure** for language model access.
  • Installation and Launch: A simplified process to kickstart AutoGen Studio.
  • Interface Navigation:
    • Build Section: Enables AI agent creation, skill definition, and workflow setup.
    • Playground Section: A dynamic platform for testing and observing AI agent behavior.
    • Gallery Section: Stores AI development sessions for future reference.
  • Powerful Python API:
    • Beneath its web interface, AutoGen Studio features a **Python API**.
    • Developers can use this API to gain detailed control over agent workflows.

MultiOn

YouTube ... Quora ...Google search ...Google News ...Bing News

Personal AI agent and life copilot that uses the browser to execute complex tasks.

  • MultiOn Browser is a web browser that uses ChatGPT and OpenAI plugins to interact with anything on the internet on your behalf
  • MultiOn is also a ChatGPT plugin that can automate tasks for you, such as posting on social media or ordering products online, or find any content on the web


CrewAI

YouTube ... Quora ...Google search ...Google News ...Bing News

CrewAI is a cutting-edge framework designed for orchestrating role-playing and autonomous AI agents. It fosters collaborative intelligence, empowering agents to work together seamlessly and tackle complex tasks. Here are some key points about CrewAI:

  • CrewAI enables AI agents to assume roles, share goals, and operate as a cohesive unit—much like a well-coordinated crew.
  • Whether you're building a smart assistant platform, an automated customer service ensemble, or a multi-agent research team, CrewAI provides the backbone for sophisticated multi-agent interactions.
  • Agent Setup Example:
  - Let's say you're creating a research team. Here's how you might set up an agent: