Skip to content

Latest commit

 

History

7 Commits

Folders and files

NameName
Last commit message
Last commit date
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 
 

Repository files navigation

Diffmedia loop

A closed-loop planner for in-vitro differentiation-media factors.

It searches a box taken from published windows (CHIR99021, IWP2, SB431542, LDN-193189), fits a small Gaussian process, and proposes the next point by expected improvement. The response it optimizes in this repository is a cartoon: Wnt on then Wnt off scores as "cardiac"; dual SMAD inhibition scores as "neural"; each punishes the other. That is not a differentiation dataset. It is a bake-off of the planner against random search at the same budget, which is the thing you can prove with no hood.

On 12 seeds and a budget of 20 evaluations, mean best score was about 0.74 vs 0.24 (cardiac cartoon) and 0.74 vs 0.16 (neural cartoon). Replace observe with a real column (qPCR, a troponin fraction, a score from brightfield-colony-qc) before anyone plates the suggestion.

Non-goals

  • Doses for a person, an animal, or a "cycle"
  • A claim that 8 µM CHIR and 4 µM IWP2 is the right cardiac protocol. Those are the peaks of the synthetic surface, on purpose, so the test can see whether search finds them. cell-protocol-compiler is where the published GiWi and dual-SMAD checklists live, including the warning that CHIR is line-dependent.
  • A foundation model of a cell. Four factors and a GP.

Run

make test
PYTHONPATH=src python3 -m medialoop.cli --objective cardiac --seeds 12
PYTHONPATH=src python3 -m medialoop.cli --objective neural --seeds 12

# Run simulation with batch acquisition
PYTHONPATH=src python3 -m medialoop.cli --objective cardiac --batch-size 4 --batch-strategy kriging_believer

Python 3.10+ and numpy.

Batch acquisition (Kriging Believer and Constant Liar)

To suggest multiple conditions per round, use --batch-size N and --batch-strategy with kriging_believer, constant_liar_min, constant_liar_max, or constant_liar_mean on medialoop or medialoop-plan. Kriging Believer uses the GP posterior mean as a temporary readout for each intermediate candidate. Constant Liar uses the minimum, maximum, or mean of observed responses as that temporary readout.

Hooking a real assay later

medialoop.loop.run takes any function from a factor dict to a float in roughly [0, 1]. A wet-lab round is the same function with a human in it: the planner prints a point inside the published box, somebody runs that well, the number comes back, the GP updates. The box is a safety rail for the cartoon, not a substitute for a protocol range check.

License

MIT.

Plan from recorded measurements (v0.2)

medialoop-plan reads a collaborator-reviewed candidate table, completed observations and pending IDs. It proposes one unused candidate, records input hashes and never calls the synthetic response function. Install with python -m pip install -e ..

medialoop-plan --candidates examples/candidates.csv \
  --observations examples/observations.csv --pending examples/pending.csv \
  --seed 0 --out artifacts/proposal.json

These example IDs, factor combinations and readouts are synthetic software fixtures, not an experimental design for a cell line. A paper reporting each factor separately does not validate their Cartesian product or timing.

  • Candidates: unique candidate_id plus all four factor columns in space.py.
  • Observations: candidate_id,response; the finite response is maximized.
  • Pending: candidate_id; completed and pending IDs cannot overlap.
  • A candidate can appear once in observations. Aggregate replicates explicitly and retain their raw data elsewhere; this GP assumes a common noise scale.
  • With fewer than two observations the selection is random and reproducible. Otherwise it uses expected improvement with fixed GP hyperparameters.
  • Without a reservation ledger, record the proposal in the pending table before asking again. This preview mode does not reserve conditions or coordinate concurrent users.
  • On completion, remove the pending ID and add its observed response.

Unknown IDs, duplicate conditions, nonfinite values, out-of-box factors and an exhausted candidate set are rejected. Output files are created exclusively. --noise is an assumed response standard deviation, not an estimated noise model. The fixed GP is a baseline and its uncertainty is not calibrated on cells.

Reserve a candidate atomically

Use one shared local SQLite ledger for all callers that need reservations:

medialoop-plan --candidates examples/candidates.csv \
  --observations examples/observations.csv --pending examples/pending.csv \
  --reserve-ledger artifacts/reservations.sqlite3 --request-id round-001 \
  --seed 0 --out artifacts/round-001.json

Selection and reservation commit in a single transaction before the proposal is returned. Concurrent processes using that ledger cannot reserve the same condition under different request IDs. Existing observed/pending CSV entries are also excluded. The command reads those CSVs and writes the ledger and proposal JSON; it never edits the CSVs. Input hashes describe the exact bytes parsed for the original proposal.

Choose a new --request-id for each new request. Retry with the same ID, candidate table, seed and noise to recover its original proposal, even after observations change. This also recovers a committed reservation after output export fails or the caller loses its connection. Retry with a new output path if the previous export is partial or contains different text; an existing exact export is accepted. A retry returns a historical proposal and does not reserve another condition. The Python equivalent is medialoop.reservations.propose_and_reserve(..., ledger_path=..., request_id=...); medialoop.planner.propose(...) remains a read-only preview.

The first successful reservation binds the ledger to the candidate CSV's exact SHA-256. Keep that table unchanged. Reservations remain recorded after their responses are added to the observations CSV, so completed conditions cannot be issued again. There is intentionally no release/reissue command. Do not delete, replace or copy the ledger to start another round of the same campaign; doing so loses coordination. If a condition was also listed in a manual pending CSV, remove that entry when adding its response to observations.

SQLite releases locks and rolls back uncommitted writes when a process exits. Callers wait up to 30 seconds for a writer before returning an error; retry with the same request ID. Use a local filesystem with reliable SQLite locking, not network shares or cloud-synced copies. All reserving callers must use this API and ledger. Manual CSV updates are outside the transaction: pause submissions while updating observations/pending files, then resume with new request IDs. Preview mode can still display already-reserved conditions because it does not read the ledger. Store the ledger alongside its recovery journal files outside Git and use SQLite-aware backups while it is active.

Acquisition bake-off (v0.3)

medialoop-bakeoff compares the current planner (expected improvement) with random search, fixed-kappa UCB, and one-draw Thompson sampling on a fixed grid inside the published windows. A fifth policy, repeat expected improvement, re-proposes the same candidate inside a batch so violations are visible. Guarded policies fantasize at the posterior mean within a batch (kriging believer) and call the cartoon only once per chosen grid point.

PYTHONPATH=src python3 -m medialoop.bakeoff --objective cardiac --seeds 6

Report simple regret against the best point on that grid, cumulative regret, and batch-uniqueness violations. This is still the synthetic surface. It is not a dose, and it is not a reason to plate a well. A neural policy is intentionally absent: a course project can add one behind the same regret and violation columns without replacing the expected-improvement baseline.

About

Closed-loop planner for in-vitro media factors. Expected improvement versus random search on a synthetic differentiation surface.

Topics

Resources

Stars

0 stars

Watchers

0 watching

Forks

Releases

Packages

Contributors

Languages