Problem statement: "SatQuery AI – An Interactive Vision-Language Assistant for Multimodal Remote Sensing Image Analysis through Text Queries" Organization: ISRO · Category: Software · Theme: Space Technology
This is a small, laptop-runnable prototype that demonstrates the core concept of SatQuery AI — a natural-language interface over remote-sensing images, driven by an agentic controller that automatically picks the right specialist tool. It is a demonstration of the architecture and orchestration idea, not a production-scale, research-grade system.
| Component | What it actually is |
|---|---|
Agentic orchestration (agent.py) |
Real, working rule-based controller that inspects the query text + image count/modality and routes to 1 of 4 tools. This is the core deliverable — genuinely executable branching logic, not a mock. |
Input validation (utils/validation.py) |
Real checks — file format, image count vs. task, modality vs. task. |
Image I/O + GeoTIFF metadata (utils/image_loader.py) |
Real. Uses Pillow for PNG/JPEG and (optionally) rasterio for GeoTIFF CRS/transform/band-count metadata. |
Single-image VQA (tools/vqa.py) |
Baseline, not a trained VLM. Rule-based HSV-color + edge-texture heuristic that estimates land-cover class percentages (vegetation/water/built-up/soil) and answers keyword-matched questions about them. |
Region grounding (tools/grounding.py) |
Simplified grounding baseline, explicitly permitted by the brief. Maps a query keyword (e.g. "water") to a fixed land-cover class, reuses the VQA heuristic mask, and draws a bounding box around the largest matching region. Not open-vocabulary / learned grounding. |
Change detection (tools/change_detection.py) |
Real classical computer-vision technique (grayscale image differencing + Otsu thresholding + connected-component cleanup) — a genuinely standard remote-sensing change-detection method, just not a deep model. |
Optical+SAR fusion (tools/optical_sar.py) |
Decision-level rule fusion baseline. Combines the optical heuristic with a SAR-intensity heuristic (dark = water, bright+textured = built-up double-bounce) using simple AND/OR logic — not a learned multi-modal fusion network. |
Every tool's docstring contains a TODO section describing exactly which
pretrained/fine-tuned model would replace that heuristic in a research-grade
version (e.g. a BigEarthNet/RSVQA-tuned VLM for VQA, GroundingDINO/CLIPSeg for
grounding, a Siamese CNN trained on LEVIR-CD for change detection, a SEN12MS
fusion network for optical+SAR). The tool interface (run(images, query) -> dict)
is intentionally the same for both the baseline and a future real model, so
swapping one in later requires no changes to agent.py or app.py.
No results are fabricated. Every number shown in the UI (percentages, bounding boxes, confidence) is computed live from the pixels you upload.
satquery/
├── app.py # Streamlit UI
├── agent.py # Agentic controller / tool registry usage
├── tools/
│ ├── __init__.py # TOOLS registry (the "model/tool registry")
│ ├── vqa.py # Single-image VQA (baseline)
│ ├── grounding.py # Text-guided region grounding (baseline)
│ ├── change_detection.py # Bi-temporal change analysis
│ └── optical_sar.py # Optical+SAR cross-modal fusion (baseline)
├── utils/
│ ├── image_loader.py # PNG/JPEG/TIFF/GeoTIFF loading + metadata
│ └── validation.py # Input/config validation
├── data/sample/
│ └── generate_samples.py # Creates SYNTHETIC demo images (see below)
├── requirements.txt
└── README.md
cd satquery
python3 -m venv venv
source venv/bin/activate # Windows: venv\Scripts\activate
pip install -r requirements.txtrasterio (used only for real GeoTIFF geo-metadata) can be tricky to install
on some systems. If it fails, just remove it from requirements.txt and
reinstall — the app still runs fine and falls back to Pillow for TIFFs; it
will simply report is_geotiff: false for those files.
No paid APIs, no GPU, no internet access, and no model-weight downloads are required to run the prototype as shipped.
Real satellite images are not bundled (to keep the repo small and avoid licensing issues). Instead, generate small synthetic placeholder images that exercise every tool end-to-end:
cd data/sample
python3 generate_samples.pyThis creates (clearly synthetic — colored blocks representing water/ vegetation/built-up/soil, not real satellite data):
single_scene.png— for VQA and grounding demoschange_before.png,change_after.png— for the change-detection demo (a new "built-up" block appears in the after image)fusion_optical.png,fusion_sar.png— matching optical/SAR pair for the cross-modal fusion demo
For an actual SIH judging demo, replace these with real chips from e.g. Sentinel-2 (optical) / Sentinel-1 (SAR) via the Copernicus Open Access Hub, or crops from the BigEarthNet dataset. The pipeline accepts PNG/JPEG or GeoTIFF/TIFF interchangeably — no code changes needed.
streamlit run app.pyThen, in the browser tab that opens:
- Upload image(s) (1 for VQA/grounding, 2 for change-detection or optical+SAR fusion).
- Optionally set the modality selector (or leave it on "Auto" and let the agent infer the task from your query + image count).
- Type a query.
- Click ANALYZE.
| # | Upload | Query | Auto-routed to |
|---|---|---|---|
| 1 | single_scene.png |
"Describe the land-cover and major objects visible in this image." | Single-Image VQA |
| 2 | single_scene.png |
"Highlight the water body referred to in the query." | Text-Guided Region Grounding |
| 3 | change_before.png + change_after.png |
"What changed between these two dates, and where did the change occur?" | Bi-Temporal Change Analysis |
| 4 | fusion_optical.png + fusion_sar.png |
"Use the optical and SAR images together to identify built-up and water-covered regions." | Optical+SAR Cross-Modal Analysis |
agent.py: run_query() implements the exact 8-step flow from the brief:
- Understand the query — lowercase + keyword scan (
agent.py's keyword lists for change/grounding/fusion intents). - Identify the task —
detect_task(query, num_images, modality_choice)picks one ofvqa | grounding | change_detection | optical_sar_analysis, using the modality selector as a strong signal first, then query keywords, then image count as a fallback. - Check the input configuration —
utils/validation.pyverifies the right number of images and (if set) the right modality are present for the detected task; mismatches raise aValidationErrorshown directly in the UI. - Select the analysis module —
tools/__init__.py: TOOLSis the model/tool registry;agent.pylooks upTOOLS[task]. - Execute the analysis — the selected tool's
run(images, query)runs and returns{answer, images, confidence, details}. - Visual evidence — the tool's
imagesdict is rendered as a row of captioned images in the UI (app.py). - Natural-language answer — the tool's
answerstring is shown. - Execution summary —
agent.pyassemblesexecution_summary(input count, modality, detected task, selected tool, confidence) andapp.pyrenders it as a table.
Because task selection is a plain function (detect_task) with a simple
signature, it can be swapped for an LLM-based classifier later (see the
TODO in agent.py's docstring) without touching the tools or the UI.
- Single-image VQA →
tools/vqa.py, routed by default for 1-image queries without grounding keywords. - A second single-image capability (region grounding) →
tools/grounding.py, implemented as the brief's suggested "clearly labeled simplified grounding demonstration." - Bi-temporal change analysis →
tools/change_detection.py, using a standard image-differencing method as the brief allows when "a sophisticated pretrained model is not practical." - Optical+SAR cross-modal analysis →
tools/optical_sar.py, using the brief's explicitly-permitted "lightweight fusion/decision approach." - Agentic orchestration →
agent.py+tools/__init__.py'sTOOLSregistry, with real branching based on query + image configuration (not a single fixed pipeline, not a chatbot). - Input validation →
utils/validation.py(image count, format, modality-vs-task checks) +utils/image_loader.py(format/metadata reading, including GeoTIFF). - Modular architecture for future model upgrades → every tool has the
same
run(images, query) -> {answer, images, confidence, details}contract and aTODOdocstring naming the exact research-grade model that would replace the current heuristic.
- Land-cover classification uses simple color/texture rules, not a trained segmentation network — it will misclassify unusual color palettes (e.g. non-standard visualizations, false-color composites) or shadows.
- Grounding only supports a fixed vocabulary (water/vegetation/built-up/soil), not arbitrary free-text referring expressions.
- Change detection assumes the two images are already roughly co-registered (same framing/extent); it does not perform geometric alignment.
- SAR handling treats any 3-channel image as if its grayscale intensity were
SAR backscatter — for a real demo, upload an actual SAR amplitude image
(or the synthetic
fusion_sar.pngprovided). - Confidence scores are fixed heuristic constants per method (not learned calibration) — they are honest "this is a rule-based method" markers, not statistically calibrated probabilities.
These are exactly the kind of "if a sophisticated model is unavailable or too large, implement a simple baseline" trade-offs the brief explicitly invites for the prototype stage.