Skip to content
View Guts1005's full-sized avatar

Highlights

  • Pro

Block or report Guts1005

Block user

Prevent this user from interacting with your repositories and sending you notifications. Learn more about blocking users.

You must be logged in to block users.

Content in all repositories owned by your account will be closed.
Maximum 250 characters. Please don’t include any personal information such as legal names or email addresses. Markdown is supported. This note will only be visible to you.
Report abuse

Contact GitHub support about this user’s behavior. Learn more about reporting abuse.

Report abuse
Guts1005/README.md

?Sharvin

Systems Software & AI Infrastructure · Distributed LLM Inference · Edge Fleet Telemetry

Integrated B.Tech + MBA (Computer Engineering) @ NMIMS MPSTME, Mumbai · Class of 2028

LinkedIn GitHub Email


sharvin@edge-node:~$ neofetch --profile
OS: Linux / Edge Fleet Architecture (6-node distributed cluster)
Host: NMIMS MPSTME, Mumbai (Integrated B.Tech + MBA Tech '28)
Kernel: Modern C++20, Python, POSIX pthreads, WebRTC
Experience: Software & Systems Intern @ Aspire Consultancy Services | Ex-Intern @ Entice Engineering
Core Focus: Distributed LLM Inference (vLLM, SGLang) & High-Throughput Edge Telemetry
Current State: Architecting offline-first streaming daemons & instrumenting AI infrastructure

❕ Executive Summary

I am a computer engineering student and systems software engineer building at the intersection of low-level systems and applied AI infrastructure.

Rather than stopping at high-level API wrappers, I work deep in the stack: contributing latency telemetry upstream to tier-1 LLM inference engines (SGLang, vLLM / FlashAttention), engineering memory-constrained C++ daemons for multi-node edge camera fleets, and shipping full-stack predictive ML platforms. My guiding ethos: if it doesn't survive network partitions and high I/O concurrency, it isn't ready for production.


😌 Upstream Systems & Open Source Contributions

A selection of upstream pull requests contributed to production AI inference engines and systems infrastructure:

Repository Pull Request / Patch Architectural Scope & Impact
sgl-project/sglang PR #38228 Distributed LLM Inference: Instrumented Prometheus latency telemetry across distributed streaming queues for high-throughput inference monitoring.
vllm-project/flash-attention PR #194 GPU Acceleration: Hoisted preprocessor directives from macro expansions for MSVC compiler conformance on NVIDIA Hopper architectures.
harsh-nod/fe2o3 PR #273 Linux Systems: Hardened Linux memfd_create file sealing against bounded EBUSY collision windows under heavy concurrent process forks.
Guts1005/gmail-oauth-mailer Repository Developer Tooling: Diagnosed and resolved async socket timeout race conditions in automated CI test harnesses; designed for modular npm packaging.

⚒ Production Systems & Work Experience

Software & Systems Engineering InternAspire Consultancy Services

(May 2026 – July 2026)

  • Distributed Edge Telemetry: Architected an offline-first distributed telemetry and video data pipeline across a 6-node edge fleet, reliably ingesting 18GB of multimodal footage and 14,750+ telemetry events across intermittent network partitions.
  • Low-Footprint C++ Daemon: Engineered an edge supervisor daemon in C++ with a <45MB RAM footprint and systemd watchdog supervision, achieving zero crash-loops and continuous automated cloud re-synchronization.
  • Sub-300ms Live Streaming: Integrated WebRTC streaming (LiveKit + FFmpeg + Next.js) and multimodal API ingestion for automated real-time inspection, dual-track recording (local & desktop), and AI snapshot comparison.

Systems Software InternEntice Engineering

(Aug 2025 – Dec 2025)

  • High-Throughput Serialization: Built event serialization routines in C and Python, hitting sub-5ms data acquisition latency under sustained high-I/O sensor throughput.
  • Hardware Integration: Developed automated hardware integration and telemetry synchronization pipelines deployed across production sensor systems.

🙈 Flagship Projects

1. ChurnIQ — Predictive Analytics & ML Pipeline Platform

Production-grade machine learning platform for customer churn analytics with automated feature transformation and real-time scoring.

  • Calibrated Ensemble Model: Engineered an end-to-end predictive pipeline analyzing high-dimensional user interaction data, achieving an 84.93% ROC-AUC with calibrated XGBoost and LightGBM models.
  • Full-Stack Architecture: Built a high-concurrency FastAPI backend with interactive React & Streamlit evaluation dashboards for lead triage and feature importance visualizers.

XGBoost LightGBM FastAPI React Streamlit


2. Smart Helmet Live (Streaming-Rpi) — Edge-to-Cloud Telemetry System

End-to-end edge-to-cloud live video and telemetry pipeline built for mission-critical remote inspection.

  • Real-Time WebRTC Streaming: Raspberry Pi camera feed hardware-encoded via FFmpeg and streamed over LiveKit WebRTC with sub-second glass-to-glass latency to a Next.js control dashboard on Vercel.
  • Edge Analytics: Integrated local and desktop synchronized recording, two-way audio channels, and edge AI snapshot comparison.

Raspberry Pi WebRTC FFmpeg Next.js Vercel

→ Streaming-Rpi Repo · → Centrix-Helmet Repo · → pi-0 Repo


📂 View Additional Engineering Projects

ShiftLeft Testing Dashboard

AI-augmented developer tooling & automated defect prediction.

  • Integrates Google Gemini API to analyze automated continuous integration test runs, predict recurring defect hotspots, and generate actionable telemetry for engineering teams.
  • Stack: Python, Google Gemini API, React, Node.js.

Gmail OAuth Mailer

Reusable, production-hardened OAuth 2.0 mail dispatch engine.

  • Abstracted complex Google OAuth 2.0 authentication flows into an npm-ready library. Hardened against asynchronous socket timeout race conditions under high-throughput queues.
  • Stack: Node.js, OAuth 2.0, Mocha/Chai.

🛠️ Technical Arsenal

Domain Technologies & Frameworks
Systems & Languages C++17/20 · C (C99/C11) · Python · TypeScript · JavaScript · SQL · Bash · Rust (Foundations)
AI Systems & LLM Infra SGLang · vLLM · FlashAttention · PyTorch · Scikit-learn · LangChain · LlamaIndex · Prometheus
Backend & Distributed FastAPI · Next.js · Node.js · React · WebRTC / LiveKit · WebSockets · Docker · Redis · PostgreSQL
Edge, Cloud & Tools Raspberry Pi · Linux Kernel (systemd/POSIX) · FFmpeg · AWS Cloud Practitioner · GitHub Actions · CMake · GDB

🔭 Research & What's on the Horizon

  • Distributed KV Cache & Memory Scheduling: Exploring PagedAttention memory layouts and chunked prefill dynamics across distributed inference instances (vLLM / SGLang).
  • Agentic Workflows via Model Context Protocol (MCP): Building multi-agent systems with deterministic tool routing, dynamic context pruning, and sandboxed execution boundaries.
  • Kernel-Level Observability: Writing eBPF probes for zero-overhead telemetry tracking on edge Linux daemons and low-latency network sockets.

🏆 Accolades & Certifications

  • 🥇 Smart India Hackathon (SIH)National Finalist (Hardware & Systems Software Track)
  • ☁️ AWS Cloud Quest: Cloud PractitionerVerified Cloud Architecture Credential
  • 📜 NPTEL Elite CertificateDesign Practices for Intelligent Product Design (IIT Kanpur)
  • 🎸 Musician & Band MemberPracticing creative discipline, live rhythm, and team dynamics

📊 GitHub Footprint & Analytics

Sharvin's GitHub Stats Sharvin's GitHub Streak

Top Languages

Let's Build Something High-Impact.

Whether you're working on distributed AI inference, edge systems, or ambitious engineering challenges — my inbox is always open.

LinkedIn Email GitHub

Pinned Loading

  1. Centrix-Helmet Centrix-Helmet Public

    Smart helmet edge computing scripts for Raspberry Pi camera capture and stream management

    Python

  2. pi-0 pi-0 Public

    Raspberry Pi zero configuration and edge device utilities for remote streaming

    Python

  3. Streaming-Rpi Streaming-Rpi Public

    Production Raspberry Pi live streaming dashboard with Next.js, LiveKit, WebRTC, and Vercel

    HTML 2

  4. sgl-project/sglang sgl-project/sglang Public

    SGLang is a high-performance serving framework for large language models and multimodal models.

    Python 35.9k 8.8k

  5. vllm-project/flash-attention vllm-project/flash-attention Public

    Forked from Dao-AILab/flash-attention

    Fast and memory-efficient exact attention

    Python 135 180

  6. harsh-nod/fe2o3 harsh-nod/fe2o3 Public

    Forked from powderluv/fe2o3

    Experimental single-source Rust GPU compiler, direct-KFD runtime, deterministic simulator, and semantic tooling for AMD GPUs.

    Rust 1