Shubh Saxena
open to work

machine learning / reinforcement learning / AI engineer

Shubh Saxena

IIT Roorkee B.Tech + M.Tech, Geophysics Minor in Economics

I build models that learn, and the systems that keep them alive in production.

Reinforcement learning, agentic AI, and applied ML, taken all the way from a research idea to a monitored, secured, deployed service. The modeling is the work; shipping it is the proof.

01 · experience

Where the shipping happened

Four roles across RL research, applied AI, and cloud engineering. Each one connects model behavior to the production system around it.

  1. Jul 2026 - Present

    RL Engineer · Polymath (YC)

    Architected the task distribution for 4 long-horizon coding-agent RL environments, specifying tool-use observation/action spaces, state-transition and reset semantics, and parameterized scenarios. Designed verifier-derived reward functions for long-horizon credit assignment, then ran distributed GRPO policy optimization with ablations over reward normalization, KL control, rollout sampling, and curriculum difficulty.

  2. Apr-Jun 2026 Cupertino, CA

    Remote MLE Intern · Decompute

    Built a memory-aware RAG system for meeting, reminder, and profile recall across temporal records, semantic profiles, and episodic chat using metadata-filtered BM25/dense retrieval and RRF, with cross-encoder reranking and CRAG correction specified for ambiguous queries. Improved resilience by 60% with reverse-proxy security, rate limiting, circuit breakers, and message queues for crash reporting; fail-closed Stripe verification maintained accurate tier-gating for 100% of users.

  3. Feb-Mar 2026 IIT Roorkee

    DevOps Intern · DoMS, IIT Roorkee

    Cut manual compliance work by 80% by building agentic RAG over 2,400 IPR pages, reaching >95% Recall@10 with MPNet, Pinecone HNSW, cross-encoder reranking, and 91.3% cache hits. Automated 70% of verification at 99.1% accuracy, then scaled retrieval and CNN inference independently on EKS using CPU demand and Redpanda consumer lag.

  4. May-Jul 2025 HCLTech

    DevOps Intern · HCLTech

    Consolidated NAT Gateways via Terraform for a 96% infra cost reduction, saving $55K+/yr. Moved Jenkins to ArgoCD GitOps with Trivy/SonarQube quality gates, cutting release cycles by 40%.

02 · systems

Models, taken to production

Every project below is a trained model or agent wrapped in the serving, evaluation, and safety machinery it needs to run for real.

swarm ai · flagship github ↗

Aegis Swarm

Autonomous drone-fleet platform: real-time telemetry, GPU vision inference, stream analytics, and LangGraph mission coordination, with an MCP policy gate on the LLM-to-actuator path so the model can never drive the hardware outside its envelope.

throughput
50-100 msg/s · 5 drones
p50 detection
~45ms · Ray Serve · YOLO · SAHI
p99 loop
~350ms · Flink · Redpanda
safety
deterministic MCP policy gate
world models · visual rl · flagship github ↗

Contrastive Latent-State Policy Learning

Reduced held-out world-model error by 28% with contrastive features and matched GRU-RSSMs over seven public-state channels. Cut live-trajectory needs by 35% using BC/IQL with two-step MOPO/MOReL/COMBO rollouts, then sustained 99.5% valid actions at sub-100ms p95 with multi-GPU recurrent PPO and four runtime gates.

prediction
28% lower held-out error
sample efficiency
35% fewer live trajectories
latency
sub-100ms p95
safety
99.5% valid actions
rl · agent eval github ↗

Agent Evaluation & Orchestration

Models Terminal-Bench tasks as an OpenEnv-compliant RL environment behind a FastAPI + WebSocket orchestration server. Cut manual evaluation overhead 4× and lifted throughput 3×.

  • OpenEnv
  • GRPO
  • FastAPI
  • Kubernetes
document ai github ↗

Intelligent Document Pipeline

A containerized OCR ensemble with selective SLM/VLM adjudication. Holds 93%+ accuracy at a sub-$0.01/document cost target across isolated Docker microservices.

  • Qwen-VL
  • YOLOv8
  • OCR
  • Docker
retrieval · rag github ↗

Cascade Intelligence

A 3-tier confidence cascade that spends the LLM only when it has to: PII scrubbing, HyDE, BM25 + FAISS, cross-encoder reranking, abstention gates, and enforced citations.

  • FAISS
  • SetFit
  • HyDE
  • Groq
03 · research

An RL idea, written up

Early work, self-published. An exploration method I built and documented while learning to take RL research end to end.

preprint Zenodo · 2026

MACE-RL: Meta-Adaptive Curiosity-Driven Exploration with Episodic Memory in RL Environments

An exploration method that adapts intrinsic curiosity rewards over time and uses episodic memory to stop agents re-exploring already-visited states. Built for sparse-reward environments where agents need a stronger signal for novelty, memory, and long-horizon discovery.

Accepted, AAAI-26 Student Abstract Program.

  • intrinsic reward
  • episodic memory
  • sparse reward
  • meta-adaptive ε
04 · stack

What I reach for

RL & World Models

  • GRPO · Recurrent PPO
  • Behavior Cloning · IQL
  • RSSMs · Model-Based RL
  • Contrastive Learning

RAG & Agents

  • MPNet · Pinecone HNSW
  • BM25 · RRF
  • Cross-Encoder Reranking
  • LangGraph · LangSmith
  • MCP · Agentic Evaluation

AI & MLOps

  • PyTorch · DDP
  • Ray Serve · FastAPI
  • MLflow · DVC
  • Redpanda · Apache Flink
  • Redis · PostgreSQL
  • Python

Cloud & Delivery

  • AWS · Kubernetes
  • Docker · Terraform
  • ArgoCD · Jenkins
  • GitHub Actions

Observability

  • OpenTelemetry · Jaeger
  • Prometheus · Grafana
  • ELK Stack

Security & Lang

  • Trivy · SonarQube
  • OWASP ZAP · Falco
  • Bash · YAML
recognition

Semi-Finalist, Pan-IIT AI/ML Hackathon (IDFC First Bank), 2026

National Rank 4, Launch Pad 2025, Pan-India Case Competition, SRCC

05 · contact

Let's put a model
into production.