Use Segment Anything 2, grounded with Florence-2, to auto-label data for use in training vision models.
-
Updated
Aug 7, 2024 - Python
Use Segment Anything 2, grounded with Florence-2, to auto-label data for use in training vision models.
GroundedSAM Base Model plugin for Autodistill
This project focuses on generating a diverse and realistic dataset for computer vision training using ChatGPT and a realistic vision image generation model. The process involves dynamically creating prompts, utilizing ChatGPT to generate image descriptions, and generating images based on those descriptions.
MCP server that lets an LLM build a custom object detector by chat: fetch images, auto-label with GroundedSAM, train YOLOv8.
Prompt based automatic annotation
Optimized ROS1 integration package for FoundationPose with Grounded SAM. Provides ROS1 wrapper nodes for real-time 6D object pose estimation with open-vocabulary segmentation. Includes timestamp synchronization and temporal filtering optimizations for Human Support Robot (HSR). FoundationPose and Grounded SAM are external dependencies.
Official implementation code for 'Would SWIR modality help for detection and segmentation in harsh weather conditions?' (ICCVW 2025)
A Cross-Frame Multimodal Retrieval Augmented Generation (CFM-RAG) for Video Intelligence. It retrieves the most relevant multimodal evidence and empowers LLMs to deliver context-rich answers.
grounded_sam_vgn_ros2_pipeline: Grounded SAM → qwen → 3D PointCloud2 projection (Isaac Sim / Gazebo) → VGN
Official implementation of KARINA (IEEE CAI 2026): LMM-generated ingredient knowledge for RGB-D nutrition estimation.
AI image masker with Grounded SAM using a Python GUI desktop app.
To associate your repository with the grounded-sam topic, visit your repo's landing page and select "manage topics."