{"slug": "alphaprotein-novo-generative-diffusion-pipeline-for-de-novo-enzyme-design", "title": "AlphaProtein Novo: generative diffusion pipeline for de novo enzyme design", "summary": "Google DeepMind released AlphaProtein Novo (AP Novo), a generative diffusion pipeline for de novo enzyme design via structural motif scaffolding that co-generates protein structures and sequences conditioned on a catalytic motif and ligand context. The package combines the diffusion model with optional LigandMPNN sequence redesign, AlphaFold 3 structure prediction, and evaluation metrics, orchestrated by run_pipeline.py, and requires pretrained generator weights (generator.bin.zst) downloaded separately from Google Cloud Storage under a separate terms-of-use agreement. Findings from the code, model parameters, or outputs must cite the bioRxiv paper \"Designing enzymes for new-to-nature chemistry and non-natural substrates with AlphaProtein Novo.", "body_md": "AlphaProtein Novo (AP Novo) is a generative diffusion pipeline for *de novo*\nenzyme design via structural motif scaffolding.\n\nThis package includes a diffusion model which co-generates protein structures\nand sequences conditioned on a catalytic motif and ligand context. This is run\nusing `run_generator.py`.\n\nTo form an end-to-end pipeline, the diffusion model is combined with optional\nsequence redesign using LigandMPNN (`run_ligandmpnn.py`), structure prediction\nusing AlphaFold 3 (`run_alphafold.py`), and evaluation metrics\n(`evaluate_design.py`). The `run_pipeline.py` script runs all stages in\nsequence.\n\nSee the [Quickstart](#quickstart) section for a basic launch command and\n[Example Design Campaigns](#example-design-campaigns) for more detailed\nexamples.\n\nAny publication that discloses findings arising from using this source code, the\nmodel parameters, or outputs produced by those should [cite](#citing-this-work)\nthe\n[Designing enzymes for new-to-nature chemistry and non-natural substrates with AlphaProtein Novo](https://www.biorxiv.org/content/10.64898/2026.10.01.756017v1)\npaper.\n\n📒 **Note: Pretrained model weights are not part of this package and must\nbe downloaded from Google Cloud Storage (see\n[Model Parameters](#model-parameters-weights) below).** Use is subject to these\n[terms of use](https://github.com/google-deepmind/alphaprotein-novo/blob/main/WEIGHTS_TERMS_OF_USE.md). Scripts look for weights under\n`./models/apnovo_generator` by default, or you can point to another directory\nusing `--apn_model_dir` (for `run_pipeline.py`), `--model_dir` (for\n`run_generator.py`), or `settings.model_dir` in your design manifest.\n\nTwo Python environments are recommended:\n\n- **Primary Environment (`alphaprotein_novo`)** : Contains JAX. Runs diffusion\ngeneration, AlphaFold 3 folding, evaluation metrics, and the pipeline\norchestrator.\n- **LigandMPNN Environment (`ligandmpnn`) [Optional]** : Contains PyTorch. Used\nto run LigandMPNN for sequence design.\n\n```\n# Create and activate a Python 3.12 environment\nuv venv --python 3.12 .venv\nsource .venv/bin/activate\n\n# Install JAX with CUDA 12 support (or CPU: uv pip install -U jax)\nuv pip install -U \"jax[cuda12]>=0.4.30\"\n\n# Install AlphaFold 3\nuv pip install git+https://github.com/google-deepmind/alphafold3.git\n\n# Install AP Novo in editable mode\ncd /path/to/alphaprotein_novo\nuv pip install -e .\n\n# Compile AlphaFold 3 chemical component data (CCD pickle)\nbuild_data\n```\n\n📒 **Tip: `build_data` compiles `ccd.pickle` and\n`chemical_component_sets.pickle` required by AlphaFold 3 to parse ligands.** If\n`components.cif` cannot be located, set the `LIBCIFPP_DATA_DIR` environment\nvariable to its directory before running `build_data`.\n\nOnly required if you plan to run optional sequence redesign with LigandMPNN:\n\n```\n# 1. Create and activate Python 3.11 environment\nuv venv --python 3.11 .venv-ligandmpnn\nsource .venv-ligandmpnn/bin/activate\n\n# 2. Clone LigandMPNN and download weights\ngit clone https://github.com/dauparas/LigandMPNN.git\ncd LigandMPNN\nbash get_model_params.sh ./model_params\n\n# 3. Install dependencies\nuv pip install -r requirements.txt\nuv pip install \"setuptools<82\"  # ProDy requires pkg_resources removed in setuptools 82+\n\n# 4. (Optional) Set environment variables. This allows you to avoid having to\n# provide the --ligandmpnn_dir and --ligandmpnn_python flags to run_pipeline.py.\nexport LIGANDMPNN_DIR=/home/<your_username>/projects/LigandMPNN  # For example.\nexport LIGANDMPNN_PYTHON=/home/<your_username>/miniconda3/envs/ligandmpnn/bin/python  # For example.\n```\n\n- **AP Novo** : Download pretrained generator weights (`generator.bin.zst` )\ninto`./models/apnovo_generator` (subject to the[AP Novo Parameters Terms of Use](https://github.com/google-deepmind/alphaprotein-novo/blob/main/WEIGHTS_TERMS_OF_USE.md) ; or configure a\ndirectory via`--apn_model_dir` ,`--model_dir` , or`settings.model_dir` in\nyour manifest):\n\n```\nmkdir -p models/apnovo_generator\nwget -P models/apnovo_generator https://storage.googleapis.com/alphaprotein_novo/generator.bin.zst\n```\n\n- **AlphaFold 3 Leaving Atom (AF3-LA)** : Download neural network weights\n(`af3_leaving_atom.bin.zst` ) into`./models/af3_la` (subject to the[AlphaFold 3 Parameters Terms of Use](https://github.com/google-deepmind/alphafold3/blob/main/WEIGHTS_TERMS_OF_USE.md) ).\nStructure prediction uses these by default, because AF3-LA is fine-tuned for\nleaving atom handling and therefore models the covalent intermediates common\nin AP Novo designs:\n\n```\nmkdir -p models/af3_la\nwget -P models/af3_la https://storage.googleapis.com/alphafold3/af3_leaving_atom.bin.zst\n```\n\nAlternatively, stock **AlphaFold 3** weights (`af3.bin.zst`) can be used,\nsubject to the same terms of use. Pass `--model_dir=./models/af3` (or\n`--af3_model_dir=./models/af3` to `run_pipeline.py`) to select them:\n\n```\nmkdir -p models/af3\nwget -P models/af3 https://storage.googleapis.com/alphafold3/af3.bin.zst\n```\n\nThe standard way to run AP Novo is via `run_pipeline.py`. It executes design\ngeneration, sequence design, folding, and metrics calculation end-to-end from a\nsingle \"manifest\" file containing configuration settings:\n\n```\n# Ensure the primary environment is active\nsource .venv/bin/activate\n\npython run_pipeline.py \\\n  --manifest=examples/kemp_eliminase/kemp_manifest.json \\\n  --output_dir=./kemp_campaign \\\n  --af3_model_dir=./models/af3_la \\\n  --ligandmpnn_dir=/path/to/LigandMPNN \\\n  --ligandmpnn_python=\"/path/to/ligandmpnn/.venv-ligandmpnn/bin/python\"\n```\n\n📒 **Tip: To quickly verify your installation without waiting for a\nfull-length diffusion trajectory, swap in\n`examples/kemp_eliminase/kemp_test_manifest.json`**, which runs the same Kemp\neliminase benchmark with reduced sampling steps.\n\nOutputs are organized into stage-specific subdirectories under `--output_dir`\n(see [Generated Artifacts & Reports](#generated-artifacts--reports) below). By\ndefault, re-running the pipeline resumes interrupted campaigns by skipping\ncompleted designs; you can also force a full rerun or execute individual stages\n(see [Modular Stage Execution](#modular-stage-execution)).\n\nCampaigns are configured using a JSON manifest describing the design problem, motif conditioning, and downstream processing.\n\nHere is an example based on `examples/kemp_eliminase/kemp_manifest.json`:\n\n```\n{\n  \"settings\": {\n    \"model_dir\": \"./models/apnovo_generator\"\n  },\n  \"defaults\": {\n    \"input_file\": \"kemp_eliminase_motif.cif\",\n    \"is_author_naming\": true,\n    \"num_sampling_steps\": 1000\n  },\n  \"designs\": [\n    {\n      \"name\": \"kemp_tight\",\n      \"motif_str\": \"A1,A2,A3|10-40,{},2-30,{},2-30,{},10-40/B1\",\n      \"motif_atoms\": \"A1:NE2,ND1 A2:OD1 A3:ND2\",\n      \"num_designs\": 200\n    }\n  ],\n  \"resequence\": {\n    \"enabled\": true,\n    \"temperature\": 0.1\n  },\n  \"folding\": {\n    \"inputs\": [\"resequenced\"],\n    \"seeds\": [230],\n    \"states\": [\n      {\"name\": \"monomer\", \"ligands\": []},\n      {\"name\": \"complex\", \"ligands\": [{\"id\": \"B\", \"ccd_code\": \"6NT\"}]}\n    ]\n  },\n  \"evaluation\": {\n    \"suite\": \"kemp_eliminase\"\n  }\n}\n```\n\nSee [below](#example-design-campaigns) for further, full-fledged examples.\n\n- **`settings`** : Run-level configuration (e.g.` model_dir` ,`output_dir` ).\nOverridden by corresponding CLI flags.\n- **`defaults`** : Default fields applied to any design job that does not\nexplicitly set them (e.g.`num_sampling_steps` ,`seed_start` ).\n- **`designs`** : List of design tasks. Each specifies the number of designs\nand the input motif (see[Design Inputs](#design-inputs) below).\n- **`resequence`** : Configures LigandMPNN sequence redesign (` enabled` ,`temperature` ). Set`\"enabled\": false` to fold the generated sequence\ndirectly without LigandMPNN.\n- **`folding`** : Specifies which structures to fold (` inputs: [\"generated\"]` or`[\"resequenced\"]` ), RNG`seeds` , and target`states` (e.g. apo monomer,\nligand complex).\n- **`evaluation`** : Sets evaluation parameters, including the enzyme\nevaluation`suite` (`\"kemp_eliminase\"` ,`\"serine_esterase\"` ,`\"dehp_esterase\"` ,`\"carbene_transfer\"` , or`\"nitrene_transfer\"` ) and an\noptional`reference_cif` override (by default, each design is scored against\nits own per-sample ground-truth motif structure).\n\n📒 **Note: Relative file paths in the manifest resolve relative to the\ndirectory containing the manifest file.**\n\nDe novo enzyme design starts with a 3D arrangement of sidechain and ligand atoms that represents a transition-state or intermediate of the target reaction. The goal of design is to generate a protein that holds this motif in the intended conformation.\n\nAt a high level, you will need to decide the following:\n\n- Input motif: what are the catalytic residues and small molecule ligands that represent your reaction? What conformation should the enzyme bind them in?\n- Full or partial residue motif: Do you want to constrain the backbone positions of all residues in the motif, or just functional groups on the sidechains?\n- Indexed or unindexed motif residues: Do you want to specify the residue indexes (primary sequence positions) of motif residues on the design ahead of generation (indexed), or have the model decide where they go (unindexed)? Unindexed conditioning can in theory find more optimal motif placements, but we find that indexed conditioning gives slightly better results in practice.\n- Novel scaffold or partial diffusion: Do you want to generate a protein from scratch, or re-use an existing protein structure as a starting point (partial diffusion)?\n\nThe design problem is specified in full using the following fields in the\n`designs` section of the manifest:\n\n- `input_file` is a path to a CIF file containing the input motif.\n- `motif_str` is a string describing which parts of the input structure make\nup the motif and how they are arranged in the output design.\n  - The string specifies a series of design chains delimited by `/` .\n  - The first design chain contains a series of comma-separated segments representing scaffolded motif or designable regions. Subsequent chains should consist of a single residue, representing a (fixed) ligand.\n  - For example, `A1,5-10,A5-6,3/B1` means \"generate a design that scaffolds\nresidue 1 of chain A of the input structure, followed by 5 to 10\nresidues of generated protein, followed by scaffolding residues 5-6 of\ninput chain A, followed by 3 residues of generated protein, with a 2nd\nchain consisting exactly of the atoms in input chain B (a ligand)\".\n  - In this example, residues 1, 5, and 6 of input chain A and 1 of chain B are \"fixed\" to their input coordinates, and the model will try to generate coordinates for the remaining residues to best preserve the position of the fixed residues or ligands.\n  - Designable length ranges are sampled before generation, e.g. for `5-10` ,\na number between 5 and 10 (inclusive) is chosen uniformly at random and\nfixed for the duration of the diffusion process.\n  - The ordering of the motif segments in `motif_str` can be sampled by\nwriting`\"A1,A5-6|{},5-10,{},3/B1\"` . This will resolve to either`A1,5-10,A5-6,3/B1` or`A5-6,5-10,A1,3/B1` with equal probability.\n- The string specifies a series of design chains delimited by \n- `seq_length` is an optional string constraining the total residue length of\nthe designed protein chain (either an exact integer like`\"150\"` or an\ninclusive range like`\"120-160\"` ). When specified, designable segment\nlengths in`motif_str` are sampled conditioned on the total chain length\nfalling within this range.\n- `unindexed_motif_residues` is an optional string containing motif residues\nfor \"unindexed\" conditioning. By default, the motif residues are at fixed,\npre-defined positions along the primary sequence (\"indexed\" conditioning).\nUnder unindexed conditioning, the model decides during generation where the\nmotif should go.\n  - This is a comma-separated list of residues or residue ranges, e.g.\n`\"A1,A2,A3\"` or`\"A120-121,A188\"` .\n  - If specified, `motif_str` must contain a single designable element (e.g.`10-250/B1` ), with no fixed residues specified other than ligand chains\n(which are always fixed).\n- This is a comma-separated list of residues or residue ranges, e.g.\n- `motif_atoms` is an optional string describing which atoms of motif residues\nare considered fixed, allowing for side chains to be generated by the model\n(and thereby be \"flexible\"). This is sometimes called \"tip atom\" or \"atomic\nmotif\" conditioning.\n  - This is a space-separated list of groups, where each group consists of a residue name and a comma-separated list of atom names.\n  - If not specified, all atoms of the residues in `motif_str` are fixed.\n  - If specified, the indicated atoms are fixed and the rest of the residue is generated by the model.\n  - For example, `\"A1:NE2,ND1 A2:OD1\"` means \"fix NE2 and ND1 of residue A1\nand OD1 of residue A2\".\n- `reseq_residues` is an optional string specifying residues on the motif to\n\"resequence\", or allow to change in amino-acid identity during the\ngeneration and resequence stages. This can be used to constrain the backbone\nposition of a residue while allowing its sidechain to vary in identity.\n- `is_author_naming` is a boolean indicating whether the motif string uses\nauthor-assigned residue numbers or PDB \"internal\" numbering. PyMOL and the\nliterature almost always use author naming. The PDB structure viewer uses\ninternal numbering.\n- `partial_diffusion_input_file` is an optional path to a CIF file containing\na starting protein structure (such as a parent design) to diversify via\npartial diffusion. Must be specified together with`partial_diffusion_num_steps` .\n- `partial_diffusion_num_steps` is an optional integer specifying how many\nreverse diffusion steps (out of`num_sampling_steps` ) to unroll when running\npartial diffusion from`partial_diffusion_input_file` . Setting this to 600\nwill lead to outputs with roughly TM-score 95 to the parent design; 900\nsteps will lead to TM-score ~80. (Total number of denoising steps is 1000.)\n\nThe `examples/` directory contains manifest files and input motif structures\n(`*.cif`) that partially reproduce the settings used for the best designs in the\n[AlphaProtein Novo paper](https://www.biorxiv.org/content/10.64898/2026.10.01.756017v1).\nProblem-specific evaluation metrics are also included in the code for each of\nthese examples, and are invoked by the manifests using the `evaluation.suite`\nfield. Note that this repo is a port of the original (Google-internal) pipeline\nused to generate the designs in the paper. Some settings (e.g. numbers of\nresequences and folding seeds) have been reduced relative to the paper so that\nexample runs can complete in a reasonable time.\n\nRun any of the example campaigns end-to-end using, for example:\n\n```\npython run_pipeline.py \\\n  --manifest=examples/kemp_eliminase/kemp_manifest.json \\\n  --output_dir=/tmp/apn_example_kemp\n```\n\nThis assumes that AP Novo and AF3 weights are in the default locations and LigandMPNN environment variables are set.\n\nExample manifests:\n\n1. \n**Kemp eliminase (`examples/kemp_eliminase/`)**\n  - Motif contains glutamate catalytic base and serine oxyanion hole H-bond\ndonor with 6-nitrobenzotriazole transition-state analog. Used for the\ndesign of `GDM_KE_1483` .\n  - Main pipeline: `kemp_manifest.json` . Unindexed conditioning variant:`kemp_unindexed_manifest.json` . Fast testing variant with 50 denoising\nsteps (will not produce good designs):`kemp_test_manifest.json` .\n  - Folds monomer (apo) and 6NT-bound states with 5 AF3 seeds each.\n  - Computes self-consistency, pocket RMSD, and Kemp eliminase catalytic geometry metrics.\n2. Motif contains glutamate catalytic base and serine oxyanion hole H-bond\ndonor with 6-nitrobenzotriazole transition-state analog. Used for the\ndesign of \n3. \n**4MU-Ac serine esterase (`examples/serine_esterase/`)**\n  - Motif is derived from cutinase 1xzm, with a Ser-His-Asp catalytic triad,\nthree oxyanion hole H-bond donors, and 4-methylumbelliferyl acetate\ntetrahedral intermediate. Used for the design of `GDM_SE_2937` .\n  - AF3 folding across all five reaction states (`monomer` ,`es` ,`complex` / TI1,`aei` , and`ti2` ), with evaluation of catalytic triad/oxyanion\nhole H-bonds and cross-state (ES/TI1) ligand RMSD and coordinate\nstandard deviation.\n4. Motif is derived from cutinase 1xzm, with a Ser-His-Asp catalytic triad,\nthree oxyanion hole H-bond donors, and 4-methylumbelliferyl acetate\ntetrahedral intermediate. Used for the design of \n5. \n**DEHPase — novel scaffold (`examples/dehp_esterase_denovo/`)**\n  - Motif is derived from kexin 1r64, with a Ser-His-Asp catalytic triad, 2\noxyanion hole H-bond donors, and bis(2-ethylhexyl) phthalate tetrahedral\nintermediate, used to design `GDM_DEHP_0176` .\n  - AF3 folding in 5 reaction states as in 4MU-Ac esterase example.\n6. Motif is derived from kexin 1r64, with a Ser-His-Asp catalytic triad, 2\noxyanion hole H-bond donors, and bis(2-ethylhexyl) phthalate tetrahedral\nintermediate, used to design \n7. \n**DEHPase — partial diffusion\n(`examples/dehp_esterase_partial_diffusion/`)**\n  - Partial-diffusion (backtracking 575 steps, out of 1000) from parent\ndesign `GDM_DEHP_0176.cif` , used to design`GDM_DEHP_0376` .\n  - Same folding and evaluation setup as in 4MU-Ac and novel-scaffold DEHPase examples above.\n8. Partial-diffusion (backtracking 575 steps, out of 1000) from parent\ndesign \n9. \n**Carbene transferase (`examples/carbene_transferase/`)**\n  - Motif with axial histidine, heme cofactor (`HEM` ), and\n(1S,2S)-cyclopropanation transition-state/product conformer, used to\ncreate`GDM_CT_0103` .\n  - Folding in 2 states: `monomer` (apo) and`complex` (`HEM` + product).\nEvaluation computes self-consistency, pocket-aligned ligand RMSD, and\ncarbene transferase heme-pocket geometry metrics.\n10. Motif with axial histidine, heme cofactor (\n11. \n**Piperidine synthase / nitrene transferase\n(`examples/nitrene_transferase/`)**\n  - Motif with axial histidine and heme-substrate cofactor complex, used to\ncreate `GDM_NT_0151` .\n  - Folding in 7 states: `monomer` (apo),`complex` ,`heme_substrate` ,`heme_piperidine_R` ,`heme_pyrrolidine_S` ,`fiveazidopentylbenzene_heme_1` ,`fiveazidopentylbenzene_heme_2` .\nEvaluation computes cross-state Fe/N self-RMSD, Fe–N distances, and\nregioselectivity based on near-attack atom distances\n(`piperidine_propensity` ).\n12. Motif with axial histidine and heme-substrate cofactor complex, used to\ncreate \n\n📒 **Note: `use_side_chain_context`**: By default, `run_ligandmpnn.py` and\n`resequence.use_side_chain_context` enable fixed-residue side-chain atom context\n(`--ligand_mpnn_use_side_chain_context=1`), which improves design outcomes. In\nthese example manifests, `\"use_side_chain_context\": false` is set to reflect the\nexact settings used in the paper, but for your own designs you should generally\nenable side-chain context.\n\n`run_pipeline.py` organizes outputs by pipeline stage:\n\n```\nkemp_campaign/\n├── pipeline_index.json         # Campaign manifest snapshot, hash, and run metadata\n├── logs/                       # Per-stage stdout and stderr logs\n│   ├── generation.log\n│   ├── resequence.log\n│   ├── folding.log\n│   └── evaluation.log\n├── metadata/                   # Run-level per-design metadata (<prefix>_metadata.json) and designs.json\n├── 01_generation/              # Stage 1: Generated structures (.cif/.pdb), motifs (_motif.cif), and sequences (.fa)\n├── 02_resequence/              # Stage 2: Resequenced structures (_reseq.cif) and sequences (.fa)\n├── 03_folded/                  # Stage 3: AlphaFold 3 predicted models, inputs, and confidence JSONs\n└── 04_eval/                    # Stage 4: Per-design evaluation JSONs and evaluation_summary.csv\n```\n\nAll files generated in Stage 1 share the prefix `<prefix>`\n(`<job>_<design_num>`, e.g. `kemp_0000`). When sequence redesign is enabled,\nStage 2 appends `_seq<NN>` (e.g. `kemp_0000_seq00`), which becomes the design\nprefix used in Stages 3 and 4.\n\n| Stage | Artifact | Description | \n|---|---|---|\n| **Metadata** | `metadata/<prefix>_metadata.json` | Run-level design metadata (fixed residues, seeds, motif paths, and folding states). | \n| **Generation** | `<prefix>.cif` ,`<prefix>.pdb` | Generated backbone structure with ligand coordinates. | \n|  | `<prefix>_motif.cif` | Ground-truth motif structure re-indexed to match the generated design. | \n|  | `<prefix>.fa` | Single-letter *de novo* amino acid sequence. | \n| **Resequence** | `<prefix>_seq<NN>_reseq.cif` | Scaffold structure with LigandMPNN-redesigned sequence. | \n|  | `<prefix>_seq<NN>.fa` ,`seqs/<prefix>.fa` | Individual and combined LigandMPNN-redesigned sequences. | \n| **Folding** | `<prefix>_<state>_folded_seed<N>.cif` | AlphaFold 3 predicted structure for the specified state and seed index `<N>` . | \n|  | `<prefix>_<state>_confidences_seed<N>.json` | AlphaFold 3 confidence scores ( `plddt` ,`ptm` ,`iptm` ,`ranking_confidence` ,`per_atom_plddt` ,`chain_pair_pde_mean` ) and the`seed` used for folding. | \n|  | `<prefix>_<state>_af3_input.json` | AlphaFold 3 folding input JSON for the specified state across all configured folding seeds. | \n| **Evaluation** | `<prefix>_evaluation.json` | Flat dictionary of metrics ( `metrics` ) for the individual design across folded states and seeds. | \n\n- `04_eval/evaluation_summary.csv` : Tabular spreadsheet written by Stage 4\nwith one row per design containing all scalar aggregated and per-seed\nmetrics across states.\n- `metadata/designs.json` : Index written by Stage 1 mapping generated designs\nto seeds, prefixes, and completion status.\n- `pipeline_index.json` : Campaign-level record written by`run_pipeline.py` with the manifest path, SHA-256 hash, completed stages, and design prefixes.\n\n`run_pipeline.py` tracks completed work and supports partial or stage-by-stage\nexecution, and each stage script can also be invoked directly.\n\n- **Resuming interrupted runs** : By default (`--resume=true` ), re-running the\npipeline command skips designs that have already finished in`--output_dir` .\n- **Forced rerun** : Pass`--noresume` to rerun all stages from scratch.\n- **Start from specific stage** : Pass`--from_stage` to resume execution\nstarting from a specific stage (`generation` ,`resequence` ,`folding` ,`evaluation` ).\n- **Run single stage** : Pass`--only_stage=folding` to run only that stage\nagainst existing outputs on disk.\n\nThe pipeline automatically scales across available hardware on a single node or across manually sharded workers:\n\n- **Multi-GPU Sharding (Stages 1 & 3)** : When multiple GPUs are visible via`CUDA_VISIBLE_DEVICES` (or detected via`nvidia-smi` ),`run_pipeline.py` —as\nwell as standalone invocations of`run_generator.py` and`run_alphafold.py` —automatically spawns one worker process per GPU\n(`--worker_id=w --num_workers=N` with`CUDA_VISIBLE_DEVICES` pinned to each\ndevice) and distributes pending designs across workers. To restrict which\nGPUs are used, set`CUDA_VISIBLE_DEVICES` (e.g.`CUDA_VISIBLE_DEVICES=0,1` ).\n- **Multi-CPU Parallelism (Stages 2 & 4)** :\n  - `run_ligandmpnn.py` runs LigandMPNN on CPU (`CUDA_VISIBLE_DEVICES=\"\"` )\nand shards pending designs across`--num_workers` parallel CPU processes\n(default:`min(16, os.cpu_count())` ), using the same worker pool to\nwrite resequenced mmCIF structures in parallel.\n  - `evaluate_design.py` evaluates folded designs in parallel across up to`min(24, os.cpu_count())` CPU worker processes when`--num_workers=1` .\n- **Manual / Multi-Node Sharding** : You can also shard Stages 1, 3, or 4\nexplicitly across separate jobs or nodes by passing`--worker_id=<0..N-1>` and`--num_workers=<N>` directly to`run_generator.py` ,`run_alphafold.py` ,\nor`evaluate_design.py` .\n\nGenerates protein backbones and initial sequences from a manifest:\n\n```\npython run_generator.py \\\n  --manifest=examples/kemp_eliminase/kemp_manifest.json \\\n  --output_dir=./samples\n```\n\n| Key Flag | Description | \n|---|---|\n| `--manifest` | (Required) Path to design manifest JSON. | \n| `--output_dir` | Directory where generated structures, sequences, and metadata are written. | \n| `--model_dir` | Model weights directory (defaults to `./models/apnovo_generator` ). | \n| `--resume` | Skip already completed designs in `--output_dir` (default:`true` ). | \n| `--return_all_data` | Include unrolled diffusion trajectory diagnostics (default: `false` ). | \n| `--worker_id` | 0-indexed worker ID for multi-GPU sharding (default: `0` ). | \n| `--num_workers` | Total number of parallel workers for sharding (default: `1` ; auto-spawned across visible GPUs when`1` ). | \n\n*Run `python run_generator.py --help` for all options.*\n\nResequences generated backbones with LigandMPNN from within the primary environment:\n\n```\npython run_ligandmpnn.py \\\n  --input_dir=./samples \\\n  --output_dir=./samples \\\n  --ligandmpnn_dir=/path/to/LigandMPNN \\\n  --python_executable=\"/path/to/ligandmpnn/.venv-ligandmpnn/bin/python\"\n```\n\n| Key Flag | Description | \n|---|---|\n| `--input_dir` | Directory holding designs to resequence (located via design metadata). | \n| `--output_dir` | Destination directory for resequenced structures and FASTA files. | \n| `--ligandmpnn_dir` | Path to cloned LigandMPNN repository. | \n| `--python_executable` | Python interpreter from the `ligandmpnn` environment. | \n| `--temperature` | Sampling temperature for sequence generation (default: `0.1` ). | \n| `--checkpoint` | Model checkpoint (default: `ligandmpnn_v_32_030_25.pt` ). | \n| `--ligand_mpnn_use_side_chain_context` | Enables side-chain context for fixed residues ( `resequence.use_side_chain_context` :`true` by default). | \n| `--num_workers` | Number of parallel CPU workers for LigandMPNN inference and structure post-processing (default: `min(16, os.cpu_count())` ). | \n\n*Run `python run_ligandmpnn.py --help` for all options.*\n\nPredicts structures using AlphaFold 3. Folding runs against the AF3-LA weights\nwith `--fix_standalone_glycans=true`, so that leaving atoms on covalent\nintermediates and unbonded glycan ligands are modelled rather than stripped:\n\n```\npython run_alphafold.py \\\n  --input_dir=./samples \\\n  --output_dir=./folded_outputs \\\n  --model_dir=./models/af3_la\n```\n\n| Key Flag | Description | \n|---|---|\n| `--input_dir` | Directory containing designs to fold. | \n| `--output_dir` | Destination directory for predicted structures and confidences. | \n| `--model_dir` | Directory containing AlphaFold 3 weights (defaults to `models/af3_la` , holding`af3_leaving_atom.bin.zst` ). | \n| `--fix_standalone_glycans` | Preserve leaving atoms on unbonded (\"standalone\") glycan ligands (default: `true` ). | \n| `--input_structure` | Structure to fold: `generated` (default) or`resequenced` . | \n| `--seed` | Random seed for AlphaFold 3 inference (default: `230` ). | \n| `--resume` | Skip designs already folded into `--output_dir` (default:`true` ). | \n| `--worker_id` | 0-indexed worker ID for multi-GPU sharding (default: `0` ). | \n| `--num_workers` | Total number of parallel workers for sharding (default: `1` ; auto-spawned across visible GPUs when`1` ). | \n\n*Run `python run_alphafold.py --help` for all options.*\n\nComputes various structural, geometric, and self-consistency metrics:\n\n```\npython evaluate_design.py \\\n  --input_dir=./folded_outputs \\\n  --output_dir=./eval_outputs\n```\n\n| Key Flag | Description | \n|---|---|\n| `--input_dir` | Directory containing folded structures and metadata. Repeatable. | \n| `--output_dir` | Directory where `<prefix>_evaluation.json` and`evaluation_summary.csv` will be written. | \n| `--eval_reference_cif` | Optional path to ground-truth motif CIF (overrides metadata). | \n| `--resume` | Reuse evaluation for designs already in `--output_dir` (default:`true` ). | \n| `--worker_id` | 0-indexed worker ID when sharding across multiple evaluation jobs (default: `0` ). | \n| `--num_workers` | Total number of shard workers when sharding across multiple evaluation jobs (default: `1` ; single-worker runs auto-parallelize across up to`min(24, os.cpu_count())` CPU processes). | \n\n*Run `python evaluate_design.py --help` for all options.*\n\n`evaluate_design.py` computes key metrics across each folded state (see\n`metrics/` for complete implementations):\n\n- \n**Self-Consistency** :\n  - \n`rmsd` :$C_\\alpha$ RMSD between the predicted model and designed scaffold (Å).\n  - \n`tm_score` : Template Modeling score between predicted and designed\nstructures (0 to 1).\n  - \n`lddt` &`gdt_ha` : Local Distance Difference Test and High-Accuracy\nGlobal Distance Test scores.\n- \n- \n**Catalytic Motif Preservation** :\n  - \n`motif_allatom_rmsd` : All-atom RMSD of catalytic motif residues with\nside-chain permutation symmetry.\n  - \n`motif_bb_aligned_allatom_rmsd` : All-atom motif RMSD after aligning\nactive-site backbones.\n- \n- \n**Ligand Pocket Geometry** :\n  - \n`mean_pocket_bb_aligned_ligand_rmsd` : Ligand RMSD after aligning pocket\nbackbone atoms. Reference-dependent ligand consistency metrics\n(`mean_pocket_bb_aligned_ligand_rmsd` ,`pocket_bb_aligned_ligand_rmsd/<ligand>` ,`mean_pocket_bb_rmsd` ,`motif_with_ligand_allatom_rmsd` , and`interface_lddt` ) are only\ncomputed for folding states whose ligand heavy-atom set (`(chain_id, res_id, res_name, atom_name)` ) and intra-residue covalent bond graph\nmatch the ligand in the generated structure; they are omitted for states\nfolded with different ligands or atom-numbering topologies (e.g. cleaved\nintermediates`aei` /`ti2` or`4MU-Ac`` es` ). When a non-covalent\nsubstrate state and covalent transition state share the same heavy-atom\nnames and bond connectivity (such as DEHP esterase`es` vs.`ti1` ),\nthese ligand consistency metrics are computed for both.\n  - \n`percent_ligand_bb_clashes_1_5` : Percentage of ligand atoms clashing\nwith protein backbone (< 1.5 Å). Computed for all ligand-containing\nfolding states.\n- \n- \n**AlphaFold 3 Confidence** :\n  - \n`plddt` : Mean per-atom predicted lDDT (0 to 100).\n  - \n`ptm` &`iptm` : Predicted TM-score and Interface predicted TM-score.\n- \n\nAny publication that discloses findings arising from using this source code, the model parameters, or outputs produced by those should cite:\n\n```\n@article{Wu2026,\n  author       = {Wu, Zachary and Abramson, Joshua and Frerix, Thomas and Chu, Alexander E. and Zhang, Ruijie K. and Schulz, Luca and Danson, Amy E. and Kwan, Tristan O. C. and Li, Wenliang K. and Kelly, Jacob and Li, Zi-Qi and Schneider, Rosalia G. and Thillaisundaram, Ashok and Patani, Harshnira and Zambaldi, Vinicius F. and Singh, Sukhdeep and La, David and Domecillo, Masy and Mora, Ariane N. and Reisenbauer, Julia C. and Zhang, Yu and Papa, Eliseo and Žemgulytė, Akvilė and Wu, Yu-Han and Žídek, Augustin and Shi, Jiaxin and Margand, Grace and Assem, Naila and Stephen, Kate and Emrich, Charlie and Liu, Peng and Colwell, Lucy and Hassabis, Demis and Fergus, Rob and Arnold, Frances H. and Kohli, Pushmeet and Wang, Jue},\n  title        = {Designing enzymes for new-to-nature chemistry and non-natural substrates with AlphaProtein Novo},\n  journal      = {bioRxiv},\n  year         = {2026},\n  elocation-id = {2026.10.01.756017},\n  doi          = {10.64898/2026.10.01.756017},\n  URL          = {https://www.biorxiv.org/content/10.64898/2026.10.01.756017v1},\n  eprint       = {https://www.biorxiv.org/content/10.64898/2026.10.01.756017v1.full.pdf}\n}\n```\n\nCopyright 2026 Google LLC\n\nAll software is licensed under the Apache License, Version 2.0 (Apache 2.0); you\nmay not use this file except in compliance with the Apache 2.0 license. You may\nobtain a copy of the Apache 2.0 license at:\n[https://www.apache.org/licenses/LICENSE-2.0](https://www.apache.org/licenses/LICENSE-2.0)\n\nThe AlphaProtein Novo Generator model parameters are made available under the\n[AlphaProtein Novo Generator Model Parameters Terms of Use](https://github.com/google-deepmind/alphaprotein-novo/blob/main/WEIGHTS_TERMS_OF_USE.md)\n(the \"APN Terms\"); you may not use these except in compliance with the Terms.\nYou may obtain a copy of the Terms at\n[https://github.com/google-deepmind/alphaprotein-novo/blob/main/WEIGHTS_TERMS_OF_USE.md](https://github.com/google-deepmind/alphaprotein-novo/blob/main/WEIGHTS_TERMS_OF_USE.md).\n\nAny pre-computed outputs of AlphaProtein Novo Generator in this repository will\nbe subject to the\n[AlphaProtein Novo Generator Outputs Terms of Use](https://github.com/google-deepmind/alphaprotein-novo/blob/main/OUTPUT_TERMS_OF_USE.md)\n(the “Output Terms”), you may not use any output except in compliance with the\nOutput Terms. You may obtain a copy of the Output Terms at\n[https://github.com/google-deepmind/alphaprotein-novo/blob/main/OUTPUT_TERMS_OF_USE.md](https://github.com/google-deepmind/alphaprotein-novo/blob/main/OUTPUT_TERMS_OF_USE.md).\n\nThe AlphaFold 3 Leaving Atom model parameters are made available under the\n[AlphaFold 3 Model Parameters Terms of Use](https://github.com/google-deepmind/alphafold3/blob/main/WEIGHTS_TERMS_OF_USE.md)\n(the \"Terms\"); you may not use these except in compliance with the Terms. You\nmay obtain a copy of the Terms at\n[https://github.com/google-deepmind/alphafold3/blob/main/WEIGHTS_TERMS_OF_USE.md](https://github.com/google-deepmind/alphafold3/blob/main/WEIGHTS_TERMS_OF_USE.md).\n\nAll other materials are licensed under the Creative Commons Attribution 4.0\nInternational License (CC-BY). You may obtain a copy of the CC-BY license at:\n[https://creativecommons.org/licenses/by/4.0/legalcode](https://creativecommons.org/licenses/by/4.0/legalcode)\n\nUnless required by applicable law or agreed to in writing, all software and materials distributed here under the Apache 2.0, the APN Terms, the Output Terms, the Terms or CC-BY licenses are distributed on an \"AS IS\" BASIS, WITHOUT WARRANTIES OR CONDITIONS OF ANY KIND, either express or implied. See the licenses for the specific language governing permissions and limitations under those licenses.\n\nYou are solely responsible for determining the appropriateness of using the software, model parameters, materials or using or distributing outputs, and assume any and all risks associated with such use or distribution and your exercise of rights and obligations under these Terms. You and anyone you share output with are solely responsible for these and their subsequent uses.\n\nOutput are predictions with varying levels of confidence and should be interpreted carefully. Use discretion before relying on, publishing, downloading or otherwise using AlphaProtein Novo Generator or AlphaFold 3 Leaving Atom.\n\nAll software, model parameters, materials and any outputs you create are for theoretical modeling only. They are not intended, validated, or approved for clinical use. You should not use the software, model parameters, materials or outputs for clinical purposes or rely on them for medical or other professional advice. Any content regarding those topics is provided for informational purposes only and is not a substitute for advice from a qualified professional.\n\nThis is not an official Google product.", "url": "https://wpnews.pro/news/alphaprotein-novo-generative-diffusion-pipeline-for-de-novo-enzyme-design", "canonical_source": "https://github.com/google-deepmind/alphaprotein-novo", "published_at": "2026-10-05 17:33:23+00:00", "updated_at": "2026-10-05 17:50:17.218066+00:00", "lang": "en", "topics": ["ai-research", "machine-learning", "generative-ai", "ai-tools", "artificial-intelligence"], "entities": ["Google DeepMind", "AlphaProtein Novo", "AlphaFold 3", "LigandMPNN", "JAX", "PyTorch", "Google Cloud Storage", "bioRxiv"], "also_reported_by": [], "alternates": {"html": "https://wpnews.pro/news/alphaprotein-novo-generative-diffusion-pipeline-for-de-novo-enzyme-design", "markdown": "https://wpnews.pro/news/alphaprotein-novo-generative-diffusion-pipeline-for-de-novo-enzyme-design.md", "text": "https://wpnews.pro/news/alphaprotein-novo-generative-diffusion-pipeline-for-de-novo-enzyme-design.txt", "jsonld": "https://wpnews.pro/news/alphaprotein-novo-generative-diffusion-pipeline-for-de-novo-enzyme-design.jsonld"}}