Type an apartment address and a date. Leasy lists every housing rule that applies, says why, and quotes the exact words of the law. When the public data cannot tell, it says unknown.
Live demo · Method note (PDF) · Judged tag: final-submission · API · Questions judges ask
Built solo for the RealPage track at Hack-Nation, 3 to 4 October 2026. The outputs being judged
are the three files in submission/, at the tag final-submission.
Note
Reviewing with a coding agent? Point it at llms.txt. It maps every guarantee to
the file and test that enforces it.
Open each link, or call the API with the same address. Each one exercises a different part of the system.
| Address | Open | What you should see | What it proves |
|---|---|---|---|
| A0001, 6238 De Longpre Ave, Los Angeles | answer page | The Rent Stabilization Ordinance (RSO) applies. The state rent cap, Civ. Code § 1947.12, shows as Replaced: superseded by the RSO. | Precedence: the stricter local rule wins, and the page says which rule governs instead. |
| A0065, 63 Bailey St | answer page | Mailed as Dorchester. Legal city Boston, MA, so Boston's rules apply. | Jurisdiction comes from the Census geocoder and 2025 city boundaries, never the mailing city. |
| A0008, 1065 Summit Ave, Jersey City, at 2027-07-01 | answer page | The NJ FAIR Act applies and is flagged for review against Jersey City's own ban. Move the date ruler back to 2026-10-01 and it shows as Upcoming, not yet effective. | Enacted is not effective. Status is computed from the effective date and the date you ask about. |
The same three checks with curl
curl "https://dchheda00--leasy-api-web.modal.run/api/addresses/A0001?as_of=2026-10-01"
curl "https://dchheda00--leasy-api-web.modal.run/api/addresses/A0065?as_of=2026-10-01"
curl "https://dchheda00--leasy-api-web.modal.run/api/addresses/A0008?as_of=2027-07-01"Each response carries the not-legal-advice notice, the jurisdiction stack and one result per rule
with its status, explanation, citation, quoted span and conflict flag. Interactive docs:
/docs.
Watch: ask a question in plain English
The chat is a router. A small model labels the question with a topic, a date and any other address. The answer itself is the engine's rows, in the engine's own words.
In weight order. These numbers come from our own scorer and a gold set drafted from organizer sources (the brief's category table and answer card, the six traps and the change tests). The official scorer and answer key are held by the judges, so treat these as our measurement, not the score.
| Area | Weight | What Leasy does | Our number | File to check |
|---|---|---|---|---|
| Extraction accuracy | 25 | A model proposes rule records from each law section. Code verifies every quote, normalises citations, parses dates from the quoted words and merges duplicates. | 86 records, 0 schema errors. Gold set: rules 25 of 25 matched, status 24 of 24, dates 10 of 10, key values 4 of 4. | submission/rules.json, eval/gold/rules_gold.json |
| Address coverage | 20 | Census coordinates, then TIGER 2025 incorporated-place polygons for the legal city. A pure engine evaluates every rule with three-valued logic. | 500 addresses at 2026-10-01, 8,255 entries: applies 6,813, unknown 627, superseded 400, pending 275, not yet effective 140. 492 of 500 placed in a city. Gold: 35 of 35 address and rule pairs. | submission/lookups.json, eval/gold/address_gold.json |
| Citations | 15 | Every quoted span is an exact slice of a stored source at stored offsets. A quote that cannot be found is retried once, then dropped. | 154 of 163 quotes found word for word. 6,127 of 6,813 applies answers (0.899) are backed by a quote from the supplied corpus. | pipeline/span_verifier.py, data/out/scoreboard.json |
| Change tracking | 15 | The engine reruns each change test at the dates it names and diffs the results. Conflicts are flagged, never resolved by a model. | T1 250, T2 90, T3 140 with 90 conflict flags, T4 110, T5 0 (an empty set is the expected answer). | submission/changes.json, engine/changes.py |
| Plain language and usability | 10 | Each answer says what the rule means for a renter, why it applies, and for an unknown, which fact is missing and where to check it. Spanish on request. | Judged on the demo. | live demo, engine/explain.py |
| Responsible design | 10 | Not-legal-advice notice on every screen and response. Unknown instead of a guess. Two other model families re-read every rule. Every model call is cached and audited. | 12 of 12 invariant tests pass. Panel agreement 0.826 (568 of 688 fields); the rest go to review. | eval/test_invariants.py, data/audit/ |
| Scalability path | 5 | A new city is data and a rerun. Typed addresses in CA, NJ and MA resolve live. The Add a law page runs a new law text through the same pipeline into a review queue. | Judged on the design. | docs/architecture.md |
Counts are from the deployed build at 2026-10-01 (/api/traps).
| Trap | What Leasy does | Count |
|---|---|---|
| Unknown is an answer | A missing building fact gives unknown, and the answer names the fact. No default year or unit count. | 212 addresses have no year built. 316 addresses have at least one unknown rule. |
| Mailing city is not the legal city | The legal city comes from TIGER 2025 polygons. The mailing city is shown, never used. | 33 addresses where they differ. 8 with no legal city, whose city rules read unknown. |
| Enacted is not in effect | Status as of a date is computed from the enactment and the effective date, kept as separate fields. | 1 rule, the NJ FAIR Act, is not yet effective at 140 addresses. It takes effect 2027-07-01. |
| Pending is not law | A pending bill is reported as pending and never applies. | 3 rules, pending at 110 addresses. |
| Struck or failed is not law | A failed measure is recorded with its status and never applies. | 2 documents classified as failed. 0 addresses show one as law. |
| The stricter rule wins | Precedence comes from exemptions extracted from the law text, not hand-coded rules. | 2 state rules superseded at 243 addresses, for example the state rent cap at A0001. |
From data/starter/dev/change_tests.json. Results are in
submission/changes.json.
| Test | What changes | Expected | Leasy |
|---|---|---|---|
| T1 | California AB 325 / SB 763 takes effect | Not yet effective on 2025-12-31, applies on 2026-01-02, at every CA address | 250 addresses |
| T2 | Hoboken and Jersey City local algorithmic bans | Each ban only inside its own city, neither in Newark | 90 addresses |
| T3 | NJ FAIR Act, enacted but not yet effective, possible preemption | Not yet effective on 2026-10-01, applies on 2027-07-02, flagged against the local bans | 140 addresses, 90 flagged |
| T4 | Massachusetts pending bills S.2983 and H.5222 | Pending, not in force, at every Boston and Cambridge address | 110 addresses |
| T5 | Massachusetts rent-control ballot question struck | No rent cap anywhere, IP 25-21 recorded as failed | 0 addresses, as expected |
Models suggest, code decides. Models classify documents and propose rule records, and nothing a model writes reaches an answer until code has checked it. Applicability, dates, status, precedence, conflicts and affected sets are all computed by deterministic code. The engine does no I/O at all, so it cannot call a model, and no model runs when you look up an address.
flowchart LR
corpus["54 corpus texts"] --> classify
links["33 linked pages,<br/>fetched once"] --> classify
classify["Classify document<br/>gpt-6-luna"] -->|law only| segment["Split into sections"]
segment --> extract["Propose rule records<br/>gpt-6-luna"]
extract --> verify{"Quote found<br/>word for word?"}
verify -->|"no: one retry, then drop"| dropped["Dropped"]
verify -->|yes| code["Dates, citations,<br/>merge duplicates"]
code --> panel["Second opinion<br/>Claude Haiku 4.5 + Gemini"]
panel --> rules[("rules.json")]
addresses["500 addresses"] --> census["Census geocoder"]
census --> tiger["TIGER 2025 polygons<br/>legal city"]
tiger --> engine["Engine<br/>three-valued logic, no I/O"]
rules --> engine
engine --> lookups[("lookups.json")]
engine --> changes[("changes.json")]
classDef model fill:#e8eaff,stroke:#4150ff,color:#1f1f1a
classDef code fill:#eef6f0,stroke:#2f6b4c,color:#1f1f1a
classDef output fill:#fff7e8,stroke:#855a10,color:#1f1f1a
class classify,extract,panel model
class segment,verify,code,census,tiger,engine code
class rules,lookups,changes output
Blue boxes call a model, green boxes are deterministic code, amber are the output files. The full
design, with the engine's logic, the date grammar, citations and confidence, is in
docs/architecture.md. Decisions and their reasons are in
docs/decisions.md.
The three judged files are in submission/, in the starter pack's formats
(data/starter/submission_templates/). Each writer checks its
file before writing, and a file that fails stops the run with nothing written.
| File | Shape | What it holds | Checked at write time |
|---|---|---|---|
submission/rules.json |
{"rules": [ ... ]} |
86 rule records, one per law: jurisdiction, level, category, status, title, plain-language requirement, key value, coverage conditions, exemptions, effective date, citation, source document and URL, the quoted span, confidence and conflict flag | Every record validated against rule_record.schema.json (JSON Schema Draft 2020-12): 0 errors |
submission/lookups.json |
{"as_of": "2026-10-01", "lookups": {"A0001": [ ... ]}} |
All 500 sample addresses, 8,255 entries. Each entry is a rule id, a result, a plain-language explanation and a conflict flag | Covers exactly the 500 ids in sample_addresses.csv |
submission/changes.json |
{"T1": {"affected_address_ids": [ ... ], "notes": "..."}, ...} |
Change tests T1 to T5. T3 also lists conflict_flag_address_ids, and each notes maps the test's rule ids to ours |
Holds exactly T1 to T5, each with an affected address list |
Values follow the schema's enums. status is one of in_force, not_yet_effective, pending or
failed; level is state or city; category is one of the six categories in the brief. In
lookups.json, result is applies (6,813), unknown (627), superseded (400), pending (275) or
not_yet_effective (140).
Example: one record from rules.json (trimmed)
{
"team_rule_id": "r-0045",
"jurisdiction": "NJ",
"level": "state",
"category": "algorithmic_rent_setting",
"status": "not_yet_effective",
"title": "Forbidding the Algorithmic Inflation of Rent (FAIR) Act",
"requirement": "The Act makes it unlawful for rental property owners to receive or contract for ...",
"key_value": null,
"coverage_conditions": null,
"exemptions": "building type in other",
"overrides": [],
"interaction": null,
"effective_date": "2027-07-01",
"citation": "P.L.2026, c.43",
"source_doc_id": "D069",
"source_url": "https://pub.njleg.state.nj.us/Bills/2026/AL26/43_.HTM",
"quoted_span": "any person to perform a coordinating function.",
"confidence": 0.92,
"conflict_flag": false,
"conflict_note": null
}status and effective_date are separate fields, so this enacted law reads not_yet_effective
until 2027-07-01. quoted_span is an exact slice of the stored source; its character offsets, the
source's SHA-256, retrieval date and source tier are in
data/out/rules_internal.json.
Example: an entry from lookups.json
{
"as_of": "2026-10-01",
"lookups": {
"A0001": [
{
"team_rule_id": "r-0012",
"result": "applies",
"explanation": "Application Screening Fee Law covers this building: address is in CA (state CA). Key figure: Actual out-of-pocket costs plus reasonable time, capped at $30 per applicant and adjustable annually by the Consumer Price Index. Not legal advice. Leasy summarises public law for information only.",
"conflict_flag": false
}
]
}
}Every explanation carries the notice. The full reasoning behind each entry (the predicates, the
facts used, any missing fact and the rule that governs instead) is in
data/out/lookup_traces.json.
Example: change test T3 from changes.json (lists trimmed)
{
"T3": {
"affected_address_ids": ["A0002", "A0003", "A0008", "..."],
"conflict_flag_address_ids": ["A0002", "A0008", "A0012", "..."],
"notes": "Mapping: NJ-ALG-01 -> r-0045. Conflict review against r-0020, r-0022 at 2027-07-02. Addresses whose result differs between 2026-10-01 and 2027-07-02."
}
}Check the files yourself. This validates every rule record against the starter schema and reads the other two files:
uv run python -c "import json, jsonschema; schema = json.load(open('data/starter/schema/rule_record.schema.json', encoding='utf-8')); rules = json.load(open('submission/rules.json', encoding='utf-8'))['rules']; [jsonschema.validate(rule, schema) for rule in rules]; lookups = json.load(open('submission/lookups.json', encoding='utf-8')); changes = json.load(open('submission/changes.json', encoding='utf-8')); print(len(rules), 'rule records valid;', len(lookups['lookups']), 'addresses at', lookups['as_of'] + ';', 'change tests', ', '.join(changes))"It prints 86 rule records valid; 500 addresses at 2026-10-01; change tests T1, T2, T3, T4, T5.
SHA-256 of the judged files:
1dbf7fa48457f2efdc5d24134a5eb8df4b660c1c0766a7d24646a663f9b0adf4 submission/rules.json
61b00c3d94bbcc1de2523f88d3e70409154b91a89d969b63f289f44e503c672f submission/lookups.json
695e0078ed234f0624a36fae9e1929b2cd0f4801a9e0d72f0683f1c8e1241a2a submission/changes.json
Everything behind the three files is committed, so any answer can be traced back to its source.
| To check | Open |
|---|---|
| The organizer's inputs: corpus, manifest, link list, addresses, schema, templates and change tests | data/starter/ |
| Full rule records with quote offsets, source SHA-256, date quotes and stage history | data/out/rules_internal.json |
| Why each address got each answer: predicates, facts, missing facts, governing rule | data/out/lookup_traces.json |
| Which city each address is in and how it was found | data/out/jurisdictions.json, data/geocode/, data/tiger/ |
| The building facts each answer used (year built, units, type, subsidy) | data/out/building_facts.json |
| How each document was classified, with the quotes behind its type and status | data/out/classifications.json, data/out/jev_check.json |
| The candidates the model proposed, and the ones dropped with a reason | data/out/candidates.json, data/out/dropped_candidates.json, data/out/consolidation_dropped.json |
| How free-text coverage was mapped to the predicate vocabulary | data/out/condition_typing.json |
| The cross-check panel's votes, and the records sent to review | data/out/crosscheck.json, data/out/review_queue.json |
| The local scorer's results and the invariant tests | data/out/scoreboard.json, eval/scorer.py |
| The gold set, drafted from organizer sources | eval/gold/rules_gold.json, eval/gold/address_gold.json, eval/gold/document_gold.json |
| The 33 link-only pages, fetched once with their status | data/supplement/ |
| Every model call (model, hashes, cache key, latency) and each pipeline step's counts | data/audit/model_calls.jsonl, data/audit/pipeline.jsonl |
| The model-call cache (tokens and cost per call) and the spend ledger | data/cache/model_calls/, data/cache/spend.jsonl |
| The accepted output of each extracted section | data/pins/sections.json |
No API keys are needed. Every model call is cached in the repository, so a rerun makes no live call and writes the same bytes.
uv sync
cp .env.example .env
uv run --env-file .env python -m pipeline.run
uv run python tools/verify.pyThe pipeline reports "live": 0 and writes files whose SHA-256 matches the hashes in
Output files. tools/verify.py runs lint, the type check, the secret and em dash
scans and all 274 tests, and ends with verify: GREEN. Setup scripts for Windows, macOS and Linux,
and fixes for common problems, are in docs/running-locally.md.
- More records than the key. 86 records against the brief's 58-rule key, mostly New Jersey tenant-guide rules and an LA anti-harassment ordinance filed under rent limits. A trimming trial lowered gold coverage, so it was not kept.
- Extraction samples vary. Earlier whole-corpus runs lost different rules on each sample. The accepted output of each section is pinned (160 sections), and a prompt change re-asks only the sections listed for it. The judges' hidden key is untested.
- Typed coverage is a reading of the law. A condition mapping can pass every source check and
still be wrong in meaning.
data/out/condition_typing.jsonlists every phrase for review. - Boston's Fair Chance policy (D010) applies on a subsidy-marked use code, a proxy that does not prove a building's funding or programme membership. Its operative date is not stated, so it stays null.
- Eight addresses have no legal city. Their geocode falls outside every incorporated place, so their city rules read unknown.
- Spanish explanations were written without a native speaker's review.
| Part | Technology |
|---|---|
| Pipeline and engine | Python 3.12, Pydantic v2, jsonschema, geopandas and shapely, managed with uv |
| API | FastAPI on Modal, one warm container |
| Web app | Next.js 16 on Vercel, MapLibre GL, React Flow |
| Geography | Census Geocoder and TIGER/Line 2025 place polygons (no key needed) |
| Role | Model | Why this one |
|---|---|---|
| Classify documents, propose rule records, verbatim retry, condition typing | gpt-6-luna |
Cheapest model that kept the hard cases right in a side-by-side test (ADR-9) |
| Second opinion on every rule | Claude Haiku 4.5 and gemini-3.5-flash-lite |
Two other model families, so their mistakes differ from the proposer's and from each other's |
| Second opinion on document type | TypeSafe Jev jev-1.13.0 |
An independent classifier for the trap documents |
| Chat router | gpt-5-nano |
Labels a question; writes no answer |
What it costs to run, at paid rates for every model including Gemini: rebuilding the whole rule set
from scratch is $0.61 (553 model calls, about $0.30 at batch rates), a changed law about $0.007,
a lookup $0 because no model runs on it, and hosting about $10 to $16 a month. One frontier model
doing the same build would cost about ten times more. Development itself, every trial and rerun
included, cost $5.59. The breakdown is in docs/costs.md.
| Read this | For |
|---|---|
docs/judges-faq.md |
The questions a reviewer is likely to ask, answered with evidence |
docs/architecture.md |
System design: components, data flow, the engine, dates, citations, confidence |
docs/api.md |
Every API endpoint, with curl examples and a Postman import |
docs/costs.md |
What it costs to run, per build, per law and per month, and what other models would cost |
docs/running-locally.md |
Install, keys, commands and setup scripts for Windows, macOS and Linux |
docs/decisions.md |
Fourteen design decisions, one paragraph each |
submission/METHOD.pdf |
The one-page method note |
| Folder | What lives there |
|---|---|
submission/ |
The three judged JSON files and the method note |
pipeline/ |
The offline pipeline, entry point pipeline/run.py |
engine/ |
The pure engine: three-valued logic, status, precedence, explanations, change tests |
api/ |
The FastAPI app and its Modal deployment |
web/ |
The Next.js web app |
eval/ |
274 tests, the invariant tests, the local scorer and the gold set |
data/ |
The starter pack, fetched pages, geocodes, boundaries, model-call cache, audit log and intermediate files |
tools/ |
The verify command |
scripts/ |
One-step setup for Windows, macOS and Linux |
Darshan Chheda, solo.
Leasy summarises public law for information only. It does not tell anyone what they may or must do, and it is not a substitute for a lawyer. Check any answer against the source it quotes.









