Skip to content

About

Galaxy (https://usegalaxy.org/) tool wrapper for BioLM.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

biolm-galaxy

Galaxy (usegalaxy.org) tool wrappers for the BioLM platform, built on the biolm CLI from biolm-sdk. Every tool shells out to biolm, so tools, notebooks, and CI all share one client, one auth model (BIOLM_TOKEN), and one backend switch (hosted platform vs. self-hosted biolm-hub).

Protocol YAML and demo fixtures are vendored under test-data/.

Tools

Tool File What it runs
BioLM Fold tools/biolm_fold.xml biolm model run esmfold predict
BioLM Encode tools/biolm_encode.xml biolm model run esm2-* encode
BioLM Predict tools/biolm_predict.xml biolm model run <model> predict
BioLM Generate tools/biolm_generate.xml biolm model run <model> generate
BioLM Protocol Runner tools/biolm_protocol_runner.xml biolm protocol run-local <protocol.yaml>

All tools share tools/macros.xml and tools/biolm_common.py.

Workflows

Workflow Pattern Protocol(s) / tools
embed_cluster.ga Protocol Runner embed_cluster
dms_landscape.ga Protocol Runner (+ collection note) dms_landscape
antibody_campaign.ga Protocol Runner antibody_campaign
library_screen.ga Protocol Runner library_screen
biosecurity_screen.ga Protocol Runner biosecurity_screen
fold_filter_embed.ga Curated-tool composition BioLM Fold → Filter (pLDDT) → BioLM Encode

Import under Galaxy's Workflows → Import. Protocol Runner workflows take an inputs_json dataset; fold_filter_embed takes FASTA. To score a Galaxy list collection with dms_landscape, collapse it to a JSON sequences array first (see the workflow annotation).

Install

Tools need biolm-sdk (which provides the biolm executable) on PATH:

pip install biolm-sdk            # curated tools
pip install "biolm-sdk[pipeline]" # + BioLM Protocol Runner (run-local)

Galaxy resolves this automatically per-tool via each <requirements> block (pip). To add these tools to a Galaxy instance, add this repo's tools/ directory (and, for legacy/advanced use, biolm_cli.xml) to your tool_conf.xml, or install via the Tool Shed / a .shed.yml-based deployment if you publish it there. For local development, planemo lint and planemo test work directly against the .xml files.

Authentication (BIOLM_TOKEN)

Every tool has a BioLM API key password field that sets BIOLM_TOKEN for that job (via each tool's <environment_variables>). Leave it blank to fall back to a BIOLM_TOKEN already present in the Galaxy job environment (e.g. set server-wide in galaxy.yml / the job's environment, or exported before starting Galaxy).

Get a token with biolm account api-key create (requires biolm account login first) or from https://biolm.ai/ui/accounts/user-api-tokens/.

Backend: hosted platform vs. self-hosted hub

Every tool exposes a Model inference backend choice:

Mode What happens How
platform (default) Model inference runs on the hosted BioLM API (biolm.ai) nothing to configure beyond BIOLM_TOKEN
hub Model inference is routed to a self-hosted biolm-hub gateway pick "hub" and supply the gateway URL (e.g. http://127.0.0.1:8000)

Internally this just sets BIOLM_BASE_API_URL for the biolm subprocess (the same variable biolm hub set writes to ~/.biolm/config.yaml, but scoped to a single job instead of mutating shared server config). Hosted protocol orchestration (biolm protocol run) always stays on the platform regardless of this switch — only model inference moves. Prefer the Protocol Runner's run-local execution (which this repo always uses) so a hub backend applies to every model task in a protocol.

The legacy biolm_cli.xml mega-tool and biolm_cli.py support the same --backend/--hub_url switch.

Protocols

Multi-step workflows are defined as YAML under test-data/protocols/ and run via BioLM Protocol Runner + biolm protocol run-local. Default protocols_root is test-data; override per-job or with BIOLM_PROTOCOLS_ROOT (see galaxy.yml).

Demo data (test-data/)

Vendored fixtures for local tool/workflow testing:

  • test-data/sequences.fasta, test-data/scaffold.pdb — inputs for the curated tools
  • test-data/<protocol_id>.inputs.json — Protocol Runner / workflow inputs
  • test-data/protocols/<protocol_id>/protocol.yaml — protocol YAML (--protocols-root test-data)

Manual smoke test (requires BIOLM_TOKEN and network access):

export BIOLM_TOKEN=...
biolm model run esmfold predict -i test-data/sequences.fasta -o /tmp/fold.json --format json
biolm protocol run-local test-data/protocols/embed_cluster/protocol.yaml \
  --input sequences='["MKTAYIAKQRQISFVKSHFSRQLEERLGLIEVQ"]' --json

Verified against the live API / known catalog gaps

The curated tools were smoke-tested end-to-end against the hosted BioLM platform (real BIOLM_TOKEN, real model calls):

  • BioLM Fold (esmfold predict), BioLM Encode (esm2-* encode), BioLM Predict (temberture-regression, biolmtox2), and BioLM Generate (antifold) all work standalone against test-data/ fixtures.
  • BioLM Protocol Runner correctly shells out to biolm protocol run-local and round-trips its --json output regardless of which protocol it's pointed at.
  • Of the bundled test-data/protocols/* fixtures, biosecurity_screen and library_screen run end-to-end via the Protocol Runner and produce real records. Protocol YAML is kept in sync with the sibling biolm-protocols catalog (biolmtox2, toxin_score ← score, mean_representation for tox encode). See that repo's KNOWN_ISSUES.md and INTEGRATION.md.

A few other bundled protocol fixtures have deeper drift between the biolm-protocols catalog and either the live model APIs or the installed biolm-sdk's local run-local runtime — these are upstream content/runtime gaps, not bugs in this repo's tool wrappers:

  • embed_cluster / antibody_campaign chain an esm2-150m encode task into a downstream depends_on task. The live esm2-150m API returns per-layer embeddings ({"embeddings": [{"layer": 30, "embedding": [...]}]}) and the installed biolm-sdk's PredictionStage stores those rows with their real layer number, but its run-local output-gating query (when no explicit layer is requested) only looks for rows with layer IS NULL — so the encode step's embeddings are computed and stored correctly, but the downstream stage sees an empty WorkingSet and gets skipped. There's no way to pin a layer from response_mapping YAML today. Workaround: run BioLM Encode and BioLM Predict/downstream steps as separate Galaxy tool steps (the pattern fold_filter_embed.ga already uses) instead of one chained protocol.
  • antibody_campaign additionally has no top-level sequences input (it's seeded from a pdb), which the installed run-local's sequence resolution rejects outright ("Could not resolve sequences from inputs").
  • dms_landscape calls esm1v-all predict with raw mutant sequences, but the live model now requires masked-marginal input (a literal <mask> token in the sequence) and rejects plain sequences.
  • inverse_fold's protein-mpnn generate response doesn't return a sequences list the way response_mapping: sequences: sequences expects (each generated design is its own top-level result with a sequence field, not nested).

None of this affects the Galaxy tool wrappers themselves — biolm_fold.py, biolm_encode.py, biolm_predict.py, and biolm_generate.py do their own response parsing (see e.g. biolm_encode.py's _extract_embedding, which already handles the embeddings/layer response shape correctly) and don't go through run-local's stage-chaining at all, so they're unaffected by the run-local limitations above. If you hit one of these gaps, prefer composing the curated tools directly in a workflow over the affected protocol, or fix the protocol YAML / upgrade biolm-sdk once the upstream issue is resolved.

Legacy / advanced: the full model/action matrix

biolm_cli.xml is an auto-generated "mega tool" exposing every model × action combination in openapi.json (260+ parameters) via a single tool, backed by biolm_cli.py. It's useful for one-off access to a model the curated tools don't cover yet, but the curated tools above are the recommended entrypoints.

Both biolm_cli.py and the generator (generate_wrapper.py) were retargeted off the deprecated biolmai SDK onto the biolm CLI (biolm-sdk), and gained the same platform/hub backend switch. To regenerate biolm_cli.xml after updating openapi.json:

python3 generate_wrapper.py

generate_tool.py is an older/unused variant that emits a differently grouped XML fragment (galaxy_model_action_inputs_grouped.xml) from the same spec; kept for reference only.

Note on planemo lint: all tools here (and the pre-existing biolm_cli.xml) use profile="21.05" idioms — <requirement type="pip">, type="password" params, and <environment_variable from_parameter="..."> — which a very recent galaxy-tool-util's strict XSD flags as errors in favor of newer <secrets>-based idioms. These are long-standing, still broadly supported Galaxy tool XML patterns; bumping to the new schema would require raising every tool's profile and is left out of scope here.

Repository layout

biolm_cli.py               legacy wrapper (invokes `biolm model run`)
biolm_cli.xml               legacy "mega tool" (auto-generated)
generate_wrapper.py         regenerates biolm_cli.xml from openapi.json
generate_tool.py            older/unused alternative generator
openapi.json                full BioLM API spec (source for the generators)
galaxy.yml                  instance user-preferences + job-env notes
tools/
  macros.xml                 shared requirements/backend/token macros
  biolm_common.py            shared subprocess/env helpers
  biolm_fold.xml / .py        ESMFold structure prediction
  biolm_encode.xml / .py      ESM2 embeddings
  biolm_predict.xml / .py     generic predict/scoring helper
  biolm_generate.xml / .py    generate/design helper
  biolm_protocol_runner.xml / .py   biolm-protocols run-local wrapper
workflows/
  *.ga                        example Galaxy workflows (see table above)
test-data/                    demo fixtures for local tool/workflow testing

About

Galaxy (https://usegalaxy.org/) tool wrapper for BioLM.

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages