My public open-source work: AI and HPC infrastructure, GPU benchmarking, inference optimization, distributed training, and governed AI. It all lives on GitHub, PyPI, and Hugging Face.

GitHub repositories

Train a model that never stops learning, no catastrophic forgetting, label-free novelty / zero-day detection, calibrated abstention.
Python · Apache-2.0
A working AI architecture built on the 2,500-year-old Vedic model of mind, one agent that learns continually without forgetting.
Python · MIT
An open, reproducible benchmark for neural-network surrogates of automotive CFD, architecture the only variable.
Python · Other
A from-scratch Mixture-of-Experts (MoE) LLM, full training pipeline, multi-domain LoRA adapters, safety guardrails, inference engine.
Python
A reproducible, vendor-neutral benchmark for MoE inference on commodity GPU clusters.
Python · Apache-2.0
Portable framework for benchmarking KV-cache, latency, and throughput of LLM inference engines (TRT-LLM vs vLLM).
Python · MIT
AI Video Analysis Platform, summarization, real-time multi-camera detection, batch processing, semantic search. Built on NVIDIA VIA + Kubernetes.
TypeScript
ns-3 simulator for AI/ML distributed-training clusters, NICs (ConnectX-4→8), Fat-tree/Dragonfly, NCCL algorithms, in-network SHARP.
Python
NCCL-based suite for GPU interconnect bottlenecks, collective-op performance, and network contention in multi-GPU training.
Python · Other
Extensible framework for evaluating and optimizing LLM inference across any quantization method, architecture, and GPU.
Shell · Apache-2.0
Production observability stack for LLM inference on Kubernetes with NVIDIA GPUs, ELK (Elasticsearch, Logstash, Kibana, Filebeat).
HTML
ARBM, a production-focused, agentic-aware evaluation framework.
Python · MIT
A reusable benchmarking tool for reasoning-first LLMs (benchmarked on NVIDIA Nemotron).
Python · Other
A framework for evaluating RAG, quality vs speed, and what actually matters.
Python
In-depth analysis of speculative decoding, 2-3× faster LLM inference without quality loss.
Python
A comprehensive framework for MoE and LLM inference performance on GPU infrastructure.
Python
MoE training benchmark on Kubernetes, EP + hybrid EP+DP; 8.77× speedup, 56% memory reduction; NCCL profiling on 4× A10.
Python
A practical guide for choosing the right LLM-training parallelism strategy.
Python
Observability for LLM inference + training on Kubernetes with NVIDIA GPUs, Prometheus, Grafana, GPU telemetry.
Python
Benchmarking suite for LLM inference frameworks on Kubernetes, NVIDIA NIM (TensorRT-LLM), vLLM, SGLang, HuggingFace TGI.
Python
Comparing LLM inference across vLLM, NVIDIA Triton, HuggingFace TGI on IBM Fusion HCI OpenShift with A100 GPUs.
Python
A complete framework for deploying and benchmarking NVIDIA cuOpt for electric-vehicle fleet route optimization.
Python · MIT
A comprehensive framework for comparing LLM inference servers on Kubernetes with NVIDIA GPUs.
Shell · MIT
GPU performance analysis & optimization for PyTorch DDP, FSDP, and DeepSpeed with NVIDIA Nsight Systems.
Shell · MIT

PyPI packages

Governed, lifelong-learning AI SDK, 50+ machine-checked invariants, provable unlearning, and a governed agent.
pip install antahkarana
Zero-dependency CLI that verifies release gates & standing invariants, ant verify --teeth.
pip install antahkarana-cli

Hugging Face models

Governed-AI SDK weights
5 days ago
Text Generation · 7B
Jun 13
36.6M params
Jun 13
Reinforcement Learning
Jun 13
Text Generation · 0.1B
May 24