Building secure, scalable and responsible AI systems from prototype to production.
Start here: Agentic AI Proof →
The proof chain is deliberately separated by function:
Hugging Face Space = experience the agent
Hugging Face Dataset = benchmark the agent
Streamlit Engineering Lab = inspect, evaluate and operate the agent
Model Quality Release Gate = decide whether a candidate model should ship
Academy · Source · Playground · Benchmark · Engineering Lab · Release Gate · Proof
I work at the intersection of AI engineering, production AI, secure systems, Edge AI, cloud/platform engineering, cybersecurity and technical leadership.
My focus is the hard part that begins after an AI prototype works: architecture, integration, deployment, evaluation, observability, governance, rollback, operational evidence and lifecycle ownership.
I build and lead systems where AI must operate reliably across enterprise platforms, industrial fleets, constrained edge devices and security-sensitive environments.
research → prototype → evaluate → secure → deploy → observe → govern → improve
My current trajectory combines hands-on engineering with platform architecture and R&D leadership:
AI Engineer → AI Architect → AI Engineering Leader → Director of AI pathway
EXPLORE • IMPLEMENT • SCALE
Agentic AI Academy is an open-source curriculum, engineering laboratory, reference architecture and professional portfolio for learning how to design, build, evaluate, secure, operate, govern and scale AI agents.
It treats agentic AI as a systems-engineering discipline rather than a collection of chatbot tutorials. The curriculum connects capability with measurable reliability, bounded autonomy, tool permissions, evaluation, observability, security, governance and safe failure behavior.
14 modules · practical engineering projects · evaluation · security · observability · governance · enterprise architecture
Hugging Face Space = experience the agent
Hugging Face Dataset = benchmark the agent
Streamlit Engineering Lab = inspect, evaluate and operate the agent
Model Quality Release Gate = convert evaluation evidence into SHIP / INVESTIGATE / HOLD
→ Academy Website
→ Proof
→ GitHub Repository
→ Agentic AI Playground
→ Evaluation & Security Benchmark
→ Agentic AI Engineering Lab
→ Model Quality Release Gate
→ Release Gate Space
→ Release Gate Dataset
curriculum → implementation → dataset → evaluation → interactive demo → release gate → production engineering
An independent reference architecture for measurable, safe, longitudinal, multimodal human-centered AI.
Human Intelligence Assurance Lab is an independent case study and public engineering reference architecture for evaluating whether emotionally aware and longitudinal AI is ready to progress toward production use.
The project focuses on measurable behavioral contracts, privacy boundaries, dependency and relationship-safety checks, uncertainty-aware interpretation, executable release policy, regression evidence and explicit SHIP / INVESTIGATE / HOLD decisions. The public stack separates interactive demonstration, datasets, model/evaluator artifacts and mutable operational evidence so the portfolio shows not only an AI concept, but also the engineering system around assurance and release readiness.
Human-Centered AI · AI Assurance · AI Safety · Longitudinal AI · Privacy · Multimodal AI · AI Evaluation · Release Gating
→ GitHub Repository
→ Live Hugging Face Space
→ Evaluation Dataset
→ Model / Evaluator Artifact
→ Operational Evidence Bucket
| Priority | Flagship | Assurance signal | Build effort | Recommendation |
|---|---|---|---|---|
| P0 | Emotional Intelligence Assurance & Release Gate | 10/10 | Medium | Build first |
| P1 | Longitudinal Memory & Relationship Safety Engine | 10/10 | Medium | Build second |
| P2 | Bio-Context Digital Twin + Trust Layer | 9/10 | High | Build after P0/P1 |
| Supporting | Custom-model architecture experiment | 7/10 | High | Research branch |
| Supporting | Polished companion UI/avatar | 5/10 | High | Don’t prioritize |
- Agentic AI : reliable tool-using agents, orchestration, RAG, MCP, evaluation, security, observability and production architecture
- Production AI Engineering : release gates, lifecycle controls, monitoring, rollback, evidence and operational reliability
- Human-Centered AI Assurance : behavioral evaluation, privacy boundaries, relationship safety, uncertainty and evidence-linked release decisions
- Secure Edge AI : trusted deployment and lifecycle control for models running across devices and industrial fleets
- AI Governance : executable policy, human approval, risk controls, attestation, audit and accountability
- Cybersecurity : device identity, attestation, trusted execution, update integrity, compromise containment and recovery
- Embedded & Trusted Systems : IoT, RISC-V, FPGA, TPM 2.0, TEE and hardware-rooted trust
Engineering trustworthy AI agents from learning to production
Open-source curriculum and engineering laboratory spanning foundations, tools, RAG, MCP, multi-agent systems, evaluation, security, production engineering, observability, governance, enterprise architecture and AI leadership.
The public stack includes a live Hugging Face agent playground, a dedicated Agentic AI evaluation and security benchmark dataset, a Streamlit engineering lab for benchmark execution and a model-quality release gate for explicit production decisions.
Agentic AI · AI Agents · LLM Engineering · RAG · MCP · AI Evaluation · AI Security · AI Governance
→ 90-sec Proof · Academy · Repository · Playground · Benchmark · Engineering Lab · Release Gate
Measurable assurance for emotionally aware, longitudinal and privacy-preserving AI
Independent reference architecture and executable evaluation system for human-centered AI. It turns behavioral expectations, dependency and sycophancy risk, privacy boundaries, wellness and crisis constraints, uncertainty and release policy into auditable evidence and explicit release decisions.
Human-Centered AI · AI Assurance · AI Safety · Privacy · Longitudinal AI · Multimodal AI · Evaluation · Release Gating
→ Repository · HF Space · Dataset · Model / Evaluator · Evidence Bucket
Evaluate. Compare. Detect regressions. Decide.
A production-oriented evaluation and release-engineering system for AI code-generation models. It compares baseline and candidate versions across helpfulness, safety, reliability, code correctness and latency; inspects failures; enforces configurable tolerances; and produces an explainable SHIP / INVESTIGATE / HOLD decision.
AI Evaluation · Model Quality · Release Engineering · AI Safety · MLOps · Hugging Face
→ Case Study · GitHub · HF Space · Dataset · Model Card
Lifecycle-First Edge AI for Industrial Fleets
A multi-partner industrial AI programme focused on turning edge-AI deployment into a governed lifecycle: optimization, secure delivery, staged rollout, qualification gates, rollback, monitoring, and operational evidence.
Edge AI · Industrial AI · AI Lifecycle · Secure Delivery · Governance · Qualification
→ life-ai.se · Demo / Platform Login
Deterministic release control for governed Edge AI
Executable governance for AI releases with fail-closed decision gates, two-person approval, risk and drift controls, signature and attestation evidence, regression qualification, auditability and MCP-based tooling.
AI Governance · Edge AI · Security · MCP · Human-in-the-Loop · Attestation
→ Repository · Live Streamlit Demo · Web Playground
Context → Action → Verification
Production-oriented AI workflow automation with typed contracts, explicit tool boundaries, business-rule verification, human approval, regression testing, CI and auditable outcomes.
Production AI · Workflow Automation · Python · Human-in-the-Loop · Testing · CI/CD
→ Repository · Live Demo
Continuous LLM Evaluation & Release Gating
Production-oriented quality assurance for LLM applications. PromptPulse turns behavioral checks into repeatable release gates across answer relevance, groundedness, reference coverage and policy compliance.
LLMOps · AI Evaluation · DeepEval · Hugging Face · GitHub Actions · Streamlit · Python
→ Repository · Live Demo
AI Engineer → AI Architect → AI Engineering Leader
A curated personal website covering production AI, Agentic AI, Edge AI, AI governance, technical leadership, research and engineering proof-of-work.
→ hendarmawan.se · GitHub Pages source
These repositories reflect broader systems, security, embedded, research and developer-tool interests:
| Repository | Focus |
|---|---|
| agentic-ai | Agentic AI Academy, curriculum, benchmark, engineering patterns and live labs |
| Human-Intelligence-Assurance-Lab | Human-centered AI assurance, longitudinal safety, privacy, multimodal evaluation and release evidence |
| model-quality-release-gate | AI model evaluation, regression detection, failure analysis and release gating |
| production-ai-automation | Production AI workflow automation and verification |
| secure-edge-ai-governance | Governed Edge AI release control and policy gates |
| PromptPulse | LLM evaluation and release gating |
| h00w.github.io | Personal portfolio and technical blog |
| autoresearch | Research-oriented engineering experiments |
| ECC-OP | Engineering / systems research repository |
| cryptobook | Cryptography-related learning and engineering material |
| yocto-v2n | Embedded Linux / Yocto engineering |
| yocto-v2n-enduser | End-user Yocto / embedded Linux work |
| Notes-and-codes | Developer productivity / notes-and-code tooling |
| claude-code-best-practice | AI-assisted software engineering practices |
| opik | LLM observability / evaluation ecosystem exploration |
| h00w | GitHub profile repository |
AI is not production-ready when the model works. It is production-ready when the system can be trusted, observed, updated, recovered and governed.
┌──────────────────────┐
│ AI / Agent │
└──────────┬───────────┘
│
┌──────────▼───────────┐
│ Evaluation Gates │
└──────────┬───────────┘
│
┌───────────────────▼───────────────────┐
│ Secure Delivery · Identity · Attest. │
└───────────────────┬───────────────────┘
│
┌──────────▼───────────┐
│ Edge / Fleet Runtime │
└──────────┬───────────┘
│
┌──────────▼───────────┐
│ Observe · Recover │
│ Evidence · Govern │
└──────────────────────┘
AI & LLM Engineering
Python · LLMs · Agentic AI · AI Agents · RAG · MCP · Prompt Engineering · AI Evaluation · LLMOps
Production AI & Platform Engineering
MLOps · Secure MLOps · Workflow Automation · GitHub Actions · Docker · Kubernetes · Observability · CI/CD
Security & Edge Systems
Edge AI · Embedded Linux · IoT · TPM 2.0 · TEE · RISC-V · FPGA · Device Attestation · Security Architecture
Leadership & Governance
AI Strategy · AI Governance · R&D Leadership · Technical Roadmaps · Research-to-Product · Industrial AI
I enjoy work that connects research depth with deployable engineering particularly secure AI infrastructure, agentic systems, production AI, industrial Edge AI, trusted computing, AI evaluation and lifecycle governance.
If you are building AI that must operate beyond the demo, I am interested in collaborations where technical depth, production reliability and responsible AI belong together.
Website · 90-sec Agentic AI Proof · Agentic AI Academy · Human Intelligence Assurance Lab · Model Release Gate · Engineering Lab · LinkedIn · GitHub · Hugging Face
Stockholm, Sweden · Hendar Mawan, PhD Eng. · @h00w

