Skip to content
splexmusPublic

About

A rag with sequences- and structure-based for protein function prediction. Using dual-rag, local-llm and graph node.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Repository files navigation

Dual-Retrieval RAG for Protein Function Research

This repository contains a Phase 1 dual-retrieval implementation: FASTA input is resolved into two independent evidence channels, normalized separately, and merged without collapsing sequence and structure evidence into one score. Both an offline fixture backend and real eggNOG-mapper + AlphaFold DB REST + local Foldseek adapters are available.

The default fixture backend requires no biological databases, downloads, or network access. Its records are synthetic and exist only to exercise schemas, provenance, agreement states, and reproducibility. It does not predict the supplied sequence's real function.

Quick start (no external data)

PYTHONPATH=src conda run -n bioinfo python -m dualrag run \
  --config config/config.yaml \
  --input tests/fixtures/queries.fasta \
  --output results/poc

The output directory contains raw channel outputs, sequence_results.tsv, structure_results.tsv, merged_results.tsv, a query manifest, a config snapshot, run metadata, and a command log.

Validation:

PYTHONPATH=src conda run -n bioinfo python -m dualrag check-tools
conda run -n bioinfo pytest -q
conda run -n bioinfo ruff check src tests

check-tools is diagnostic: missing executables are reported, but fixture mode remains fully usable. External mode fails closed until real database versions replace the UNVERIFIED placeholders.

Accepted structure accessions are explicit UniProt-style FASTA identifiers such as P69905 or sp|P69905|HBA_HUMAN. Other headers are recorded as unresolved and retain the sequence channel without making an AlphaFold request.

Scientific boundaries

  • Fixture annotations are synthetic test data, never biological evidence.
  • Scores are deterministic PoC normalization features, not calibrated confidence.
  • CONFLICT is retained as evidence; it is never averaged away.
  • Missing metrics are blank TSV cells and explicit None values internally.
  • No LLM or literature layer is present in this milestone.
  • An incomplete target database is not a research-complete search space.

See ARCHITECTURE.md, DATA.md, and docs/POC.md for contracts and extension points.

For the full external-HDD database setup, including the writable-mount preflight, eggNOG installation, Foldseek PDB download, provenance finalization, and validation, see Phase 1 data installation. The current production profile is config/4tb-hdd.yaml.

After the core eggNOG and Foldseek databases are installed, preview or run the resumable full-data downloader with:

scripts/download_phase1_full_data.sh --data-root "$DUALRAG_DATA_ROOT" --dry-run
scripts/download_phase1_full_data.sh --data-root "$DUALRAG_DATA_ROOT" --yes

About

A rag with sequences- and structure-based for protein function prediction. Using dual-rag, local-llm and graph node.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages