LAI bundle¶
The local-ancestry-inference (LAI) bundle powers the optional Tier-2 ancestry "chromosome painting." It is the heaviest artifact to build — a multi-phase, multi-hour pipeline best run on a compute cluster.
The authoritative procedure lives in the repo
The complete, step-by-step operator runbook — including exact host conventions, accuracy
sign-off gates, and the publish/rollback steps — is
docs/lai-bundle-release-runbook.md,
with its environment lock in
docs/lai-bundle-release-runbook-env.lock.yaml.
This page is an orientation summary.
The Gnomix training runtime has a separate, stronger contract: its reviewed
source specification
generates an immutable linux-64
conda-lock artifact.
Phase 05 records that lock's raw SHA-256 in every chromosome provenance record,
and Phase 07 publishes the lock in the bundle. Changing any lock byte requires
a fresh environment name, retraining all 22 chromosomes, and publishing a new
bundle generation; historical environment-free models are never backfilled.
What's in the bundle¶
A tarball containing the per-chromosome local-ancestry models, the phasing reference panel
(Beagle), genetic maps, the liftover chain, and a metadata.json carrying the bundle version
and tool versions. It is published as a GitHub release asset and pinned in
bundles/manifest.json.
The build pipeline¶
The build is a sequence of idempotent phases under
scripts/lai_bundle_v2/,
orchestrated by run_rebuild.sh (sequential) or run_rebuild_slurm.sh (cluster):
| Phase | Does |
|---|---|
| 01 | Download the phased reference panel (gnomAD HGDP + 1KG) and genetic maps |
| 02 | Lift the union catalog GRCh37 → GRCh38 and build the sites/regions files |
| 03 | Subset the reference panel to the union sites |
| 04 | ADMIXTURE-filter to a single-ancestry reference sample map |
| 05 | Train the per-chromosome ancestry models |
| 06 | Validate phasing and local-ancestry accuracy against held-out samples |
| 07 | Assemble the tarball, write metadata.json, and emit checksums |
All paths and parameters are set via
scripts/lai_bundle_v2/env.sh
so the build is cluster-portable.
Running on a cluster (SLURM)¶
Heavy phases run as a small SLURM DAG: a resumable download job (phase 01), a prep
job (phases 02–04), a train job that parallelises phase 05 across the 22 autosomes
as a job array, and a finish job (phases 06–07). Each downstream job uses an
afterok dependency, so a clean workdir cannot reach training before Phase 01 has
produced its panel and genetic-map inputs. The standard pattern is:
rsyncthescripts/lai_bundle_v2/directory to the cluster working directory.- On the login node, activate the build environment and submit
run_rebuild_slurm.sh. - Build in a shared cluster workdir visible at the same path from every SLURM node; copy only the final tarball back.
Partition, CPU, and array sizing are tunable via environment variables (see the runbook).
Validation gate¶
The bundle is only published once it clears the accuracy sign-off in the runbook (mean
per-window local-ancestry accuracy and phasing switch-error thresholds, plus held-out
per-superpopulation checks). Then it follows the same draft → verify → publish flow as the
other bundles, using the bundle-release.yml workflow with bundle_key=lai_bundle.