Protein Folding & Discovery
YouTube ... Quora ...Google search ...Google News ...Bing News
- Life Sciences ... Evolutionary ... Bioinformatics ... Proteins ... Genome ... Genealogy ... Pharmaceuticals ... Chemistry
- Symbiotic Intelligence ... Bio-inspired Computing ... Neuroscience ... Connecting Brains ... Nanobots ... Molecular ... Neuromorphic
- Case Studies
- COVID-19
- Service Capabilities
- Architectures for AI ... Enterprise Architecture (EA) ... Enterprise Portfolio Management (EPM) ... Architecture and Interior Design
- (Deep) Residual Network (DRN) - ResNet
- RoseTTAFold: Accurate protein structure prediction accessible to all | Institute for Protein Design - University of Washington
- AI Reveals Previously Unknown Biology – We Might Not Know Half of What’s in Our Cells | University of California-San Diego
- Artificial intelligence powers protein-folding predictions | Michael Eisenstein - Nature
- Artificial intelligence has worked out the structures of 200 million proteins (that’s practically all of them) | Stephanie Pappas - Livescience
- Accurate structure prediction of biomolecular interactions with AlphaFold 3 | Abramson et al. - Nature ... A breakthrough in predicting the structure and interactions of all life's molecules.
Proteins are made up of hundreds or thousands of amino acids, and these amino acid sequences specify the protein's structure and function. But understanding just how to build these sequences to create novel proteins has been challenging. Past work has resulted in methods that can specify structure, but function has been more elusive. Machine learning reveals recipe for building artificial proteins | Emily Ayshford - University of Chicago - PhysOrg
Molecular AI & Bio-inspired Computing
The synthesis of Molecular Artificial Intelligence and Bio-inspired Computing has fundamentally altered the landscape of biological engineering. By treating protein folding as a computational optimization problem, researchers have moved beyond mere prediction to *de novo* protein design. This approach leverages deep learning architectures—inspired by the evolutionary processes of natural selection—to generate entirely new protein sequences that fold into stable, functional 3D structures. These AI models act as "generative architects," allowing scientists to specify desired functional constraints (such as binding affinity or catalytic activity) and derive the necessary amino acid sequences to achieve them, effectively reversing the traditional biological discovery pipeline.
What is the protein folding problem?
Proteins are large, complex molecules essential to all of life. Nearly every function that our body performs—contracting muscles, sensing light, or turning food into energy—relies on proteins, and how they move and change. What any given protein can do depends on its unique 3D structure. For example, antibody proteins utilised by our immune systems are ‘Y-shaped’, and form unique hooks. By latching on to viruses and bacteria, these antibody proteins are able to detect and tag disease-causing microorganisms for elimination. Collagen proteins are shaped like cords, which transmit tension between cartilage, ligaments, bones, and skin. Other types of proteins include Cas9, which, using CRISPR sequences as a guide, act like scissors to cut and paste sections of DNA; antifreeze proteins, whose 3D structure allows them to bind to ice crystals and prevent organisms from freezing; and ribosomes, which act like a programmed assembly line, helping to build proteins themselves. The recipes for those proteins—called genes—are encoded in our DNA. An error in the genetic recipe may result in a malformed protein, which could result in disease or death for an organism. Many diseases, therefore, are fundamentally linked to proteins. But just because you know the genetic recipe for a protein doesn’t mean you automatically know its shape. Proteins are comprised of chains of amino acids (also referred to as amino acid residues). But DNA only contains information about the sequence of amino acids–not how they fold into shape. The bigger the protein, the more difficult it is to model, because there are more interactions between amino acids to take into account. AlphaFold: Using AI for scientific discovery | A. Senior, J. Jumper, D. Hassabis, and P. Kohli - DeepMind
Google DeepMind AlphaFold
Youtube search... ...Google search
- Google DeepMind AlphaFold
- Google DeepMind AlphaGo Zero ... * Google DeepMind AlphaStar
- 13th Community Wide Experiment on the Critical Assessment of Techniques for Protein Structure Prediction
- Protein Data Bank | Wikipedia
- AlphaFold: a solution to a 50-year-old grand challenge in biology | DeepMind
- AI has cracked a problem that stumped biologists for 50 years. It’s a huge deal. | Sigal Samuel - Vox
- DeepMind and European Molecular Biology Laboratory (EMBL) release the most complete database of predicted 3D structures of human proteins | SciTechDaily Partners use AlphaFold, the AI system recognized last year as a solution to the protein structure prediction problem, to release more than 350,000 protein structure predictions including the entire human proteome to the scientific community.
- Highly accurate protein structure prediction with AlphaFold | J. Jumper, R. Evans, D. Hassabis - Nature
- FastFold: Reducing AlphaFold Training Time from 11 Days to 67 Hours | S. Cheng, R. Wu, Z. Yu, B. Li, X. Zhang, J. Peng, Y. You
- Google DeepMind Introduces AlphaFold 3: A Revolutionary AI Model that can Predict the Structure and Interactions of All Life’s Molecules with Unprecedented Accuracy | Asif Razzaq - MarketTechPost
AlphaFold 3 & Biomolecular Interactions
AlphaFold 3 represents a significant leap beyond its predecessors by moving from simple protein structure prediction to the comprehensive modeling of biomolecular interactions. While earlier versions focused on the folding of individual protein chains, AlphaFold 3 utilizes a diffusion-based architecture to predict the structures of complexes involving proteins, DNA, RNA, and small-molecule ligands. This capability is critical for understanding how proteins interact with drugs, cofactors, and other cellular components, providing a holistic view of the molecular machinery of life. By accurately modeling protein-ligand interactions, AlphaFold 3 enables researchers to simulate how potential therapeutic compounds bind to target proteins, significantly accelerating the early stages of drug discovery.
AlphaFold Demo
Google DeepMind and Isomorphic Labs have announced that scientists will have free access to most features of their newly launched research tool, the AlphaFold Server.
To use the AlphaFold server demo, you need to start by accessing the AlphaFold server website: https://AlphaFoldServer.com
Once there, you may need to create an account or log in if you haven't already.
Next, prepare your input data by creating a FASTA file that contains the amino acid sequence of the protein you wish to analyze. The FASTA format is a simple text-based format for representing nucleotide or peptide sequences, where each sequence is preceded by a description line starting with a ">" character. Make sure your FASTA file is correctly formatted and accurately represents the protein sequence.
Once your FASTA file is ready, go to the web interface of the AlphaFold server. Here, you will find an option to upload your FASTA file, usually through a clear "Upload" button. After uploading the file, you may need to provide additional details about the protein or adjust specific parameters for the prediction. This step ensures that the AI model has all the necessary information to process your job accurately.
After submitting the FASTA file and any necessary information, you can start the prediction job by clicking the button to begin the process, typically labeled "Submit," "Run," or something similar. The processing time can vary depending on the complexity of the protein and the current server load. You might see a progress bar or receive a notification once the job is complete.
When the prediction is finished, the server will provide an overview of the protein structure as a 3D rendering. This rendering is usually interactive, allowing you to manipulate and explore the 3D model directly within your web browser. Additionally, there should be options to download the predicted structure in various formats, such as PDB, for further analysis or use in other bioinformatics tools.
Finally, you can utilize the predicted structure for your research purposes. This might involve conducting further computational analysis, planning laboratory experiments, or integrating the structure into scholarly work. The streamlined process offered by the AlphaFold 3 server demo makes it significantly easier for researchers to obtain high-quality protein structure predictions and apply these insights to their scientific endeavors.
In summary, using the AlphaFold 3 server involves accessing the web interface, preparing and uploading a FASTA file, submitting the job, waiting for the prediction to complete, and then reviewing and utilizing the results. This user-friendly approach enables researchers to conduct sophisticated protein structure analysis with minimal effort, supporting a wide range of biological and biomedical research projects.
Baker Lab
- Gaming ... Game-Based Learning (GBL) ... Security ... Generative AI ... Games - Metaverse ... Quantum ... Game Theory ... Design
- Baker Lab
David Baker is a Nobel Prize-winning biochemist known for pioneering computational protein design, primarily through the Rosetta software suite he developed, which predicts protein structures and designs novel ones, enabling new medicines and enzymes by creating proteins from scratch to solve problems nature hasn't. His lab's innovations include:
- Rosetta Software: The foundational computational platform for protein structure prediction and design, initially developed in his lab, now a large collaborative community (Rosetta Commons).
- Protein Design: Creating new proteins with desired functions, moving beyond nature's limitations for medicine, catalysts, and sustainable solutions, earning him the 2024 Nobel Prize in Chemistry.
- RoseTTAFold & RFdiffusion: AI tools that use deep learning to rapidly and accurately predict protein structures, democratizing structure determination. Recent 2024-2025 advancements have introduced multi-state protein modeling, allowing researchers to simulate conformational changes and protein flexibility, which are now integrated into open-source bioinformatics workflows to support high-throughput drug screening.
- Foldit: A game that leverages human puzzle-solving skills for protein folding challenges, engaging the public in scientific research.
Impacts:
- Revolutionized Biology: His work transformed protein design from fringe science to a powerful engineering discipline, allowing creation of molecules for therapies (like antivirals) and materials.
- Open Science: Fosters collaboration through the Rosetta Commons and open-source tools, accelerating discovery.
Pharmaceutical Applications
The integration of generative AI and advanced protein folding models is fundamentally transforming the pharmaceutical industry, particularly in the lead optimization phase of drug discovery. By accurately predicting how small-molecule ligands bind to target proteins, AI platforms like AlphaFold 3 and RoseTTAFold allow researchers to virtually screen millions of compounds with high precision, significantly reducing the time and cost required to identify promising drug candidates. This "in silico" approach enables the design of more specific, potent, and safer therapeutics, as researchers can model potential off-target effects and binding dynamics before ever entering the laboratory. This shift from trial-and-error experimentation to AI-driven rational design is shortening development cycles and opening new avenues for treating complex diseases that were previously considered "undruggable."
Future Directions
The field is currently transitioning from static protein structure prediction to the dynamic simulation of protein-protein interactions within complex cellular environments. Future research is focused on modeling the "interactome"—the complete network of molecular interactions within a cell—to understand how proteins function in real-time under varying physiological conditions. By incorporating temporal dynamics and environmental variables, these next-generation simulations aim to provide a predictive model of cellular behavior, which will be essential for developing personalized medicine and understanding the systemic impact of therapeutic interventions at the molecular level.
FASTA & FASTQ
FASTA and FASTQ are two widely used file formats in bioinformatics for storing nucleotide or peptide sequences. Here’s a detailed explanation of each:
FASTA is used to represent nucleotide sequences or protein sequences. The structure of a FASTA file includes a header line that begins with a '>' character followed by a description (e.g., sequence name or identifier), and sequence lines that contain the actual nucleotide or protein sequence in a single-letter code, which can span multiple lines.