This repository contains a Phase 1 dual-retrieval implementation: FASTA input is resolved into two independent evidence channels, normalized separately, and merged without collapsing sequence and structure evidence into one score. Both an offline fixture backend and real eggNOG-mapper + AlphaFold DB REST + local Foldseek adapters are available.
The default fixture backend requires no biological databases, downloads, or network
access. Its records are synthetic and exist only to exercise schemas, provenance,
agreement states, and reproducibility. It does not predict the supplied sequence's
real function.
PYTHONPATH=src conda run -n bioinfo python -m dualrag run \
--config config/config.yaml \
--input tests/fixtures/queries.fasta \
--output results/pocThe output directory contains raw channel outputs, sequence_results.tsv,
structure_results.tsv, merged_results.tsv, a query manifest, a config snapshot,
run metadata, and a command log.
Validation:
PYTHONPATH=src conda run -n bioinfo python -m dualrag check-tools
conda run -n bioinfo pytest -q
conda run -n bioinfo ruff check src testscheck-tools is diagnostic: missing executables are reported, but fixture mode remains
fully usable. External mode fails closed until real database versions replace the
UNVERIFIED placeholders.
Accepted structure accessions are explicit UniProt-style FASTA identifiers such as
P69905 or sp|P69905|HBA_HUMAN. Other headers are recorded as unresolved and retain
the sequence channel without making an AlphaFold request.
- Fixture annotations are synthetic test data, never biological evidence.
- Scores are deterministic PoC normalization features, not calibrated confidence.
CONFLICTis retained as evidence; it is never averaged away.- Missing metrics are blank TSV cells and explicit
Nonevalues internally. - No LLM or literature layer is present in this milestone.
- An incomplete target database is not a research-complete search space.
See ARCHITECTURE.md, DATA.md, and docs/POC.md for contracts and extension points.
For the full external-HDD database setup, including the writable-mount preflight,
eggNOG installation, Foldseek PDB download, provenance finalization, and validation,
see Phase 1 data installation. The current production
profile is config/4tb-hdd.yaml.
After the core eggNOG and Foldseek databases are installed, preview or run the resumable full-data downloader with:
scripts/download_phase1_full_data.sh --data-root "$DUALRAG_DATA_ROOT" --dry-run
scripts/download_phase1_full_data.sh --data-root "$DUALRAG_DATA_ROOT" --yes