View portfolio as
Education
M.S. in Artificial Intelligence
University of California, Santa Cruz
B.E. in Computer Science
SSN College of Engineering
Experience
AI/ML Engineer (Founding Engineer)
Elas
Built 160-scenario eval framework with golden replay harness for agentic correctness (routing, tool selection, code gen)
Designed routing across 10+ specialized agents with quality gates and regression feedback before deploy
Authored 168-page developer docs and automated release pipelines for the agent platform
Multi-AgentAgent EvalsGolden ReplayMCPArgo CI/CD
AI Research Engineer
Samsung Research Capstone
Autonomous GUI agent with event capture/replay to verify long-horizon agentic workflows
Validated 1.5-bit quantized 32B Qwen on 2x RTX 3070 with systematic regression testing
Built validation pipelines checking tool selection, scheduling, and task completion
GUI AgentsQuantizationQwenvLLMAgent Evals
Software Development Engineer
Citi Corp
Automated data validation pipelines (Hadoop, Spark) — 25% faster processing
Event-driven services processing millions of events daily
HadoopSparkData PipelinesKafka
NLP Engineering Intern
MultiOn (Stanford Startup)
Merged LangChain PR #12392; +40% agentic accuracy with retrieval + vector DBs
Open source: Tunix #1552, Ivy #13280 (merged)
LangChainRAGVector SearchOpen Source
Research
SemEval 2025 Task 4 Winner - Best Paper Award
Graduate Researcher
Chenguang Wang's Lab — UC Santa Cruz
Multi-Agent Interpretability, Agentic ML
Contributing to Google Tunix (PR #1552), Ivy (merged), massgen, and rllm
Collaboration with AutoGen creator Chi Wang (DeepMind) on consensus-based multi-agent frameworks
SemEval 2025 Task 4 Winner — Best Paper Award (hallucination detection); 4 published ML papers
Multi-Agent SystemsInterpretabilityGoogle TunixPyTorchmassgenrllm
Research Assistant
HPC Lab — Anna University (SSNCE)
Multimodal ML, Robotics
Built CLIP-based multimodal routing with automated validation (0.90 confidence)
Designed 6-DOF robotic arm with test-driven optimization (73% accuracy)
Published benchmark on HuggingFace; multiple ML papers (FIRE, journals)
CLIPPyTorchRoboticsHuggingFaceXGBoost
Projects
LLM Evaluation Framework
Built automated eval pipeline for RAG systems measuring faithfulness, relevance, and hallucination rate across model versions; supports regression detection and multi-model comparison.
AI Agent Framework
End-to-end agentic framework with SQS-backed task pipelines, Hybrid RAG with pgvector, MCP tool orchestration, and comprehensive agent evaluation suite.
Multimodal RAG System
Advanced retrieval-augmented generation system combining text, image, and audio modalities for enhanced knowledge retrieval.
Technical Skills
Ai And Ml LLMs · Multi-Agent Orchestration · Agent Evals · RAG · PyTorch · vLLM · GRPO · DDP · Quantization · NLP
Testing And Validation Eval Pipelines · DeepEval · Golden Replay · Integration Testing · Regression Testing · Langfuse
Languages Python · C++ · Go · SQL · Bash
Tools And Infra LangChain · MCP Servers · HuggingFace · Docker · Kubernetes · AWS (ECS, SQS)
Publications
A Novel Dataset for Fake News in Tamil
Abusive and Threatening Language Detection in Native Urdu Tweets
Detecting Malicious IoT Traffic Using Machine Learning Techniques
LeSS Agile Projects: A Machine Learning-Driven Empirical Model
Recognition