faster autoregressive
generation
Measured benchmark · Autoregressive generation
AI / ML ENGINEER
Open to opportunitiesBetter systems.
Backed by evidence.
I build, profile, and improve LLM systems. From GPU inference to evaluation pipelines, I turn bottlenecks into measurable progress.
faster autoregressive
generation
Measured benchmark · Autoregressive generation
01 / Selected work
Real systems. Controlled experiments.
The gains, the tradeoffs, and what I learned.
Measured benchmark · Latency tradeoffs documented
A reproducible FP8 serving benchmark that cut model weights by 42.8% and increased output throughput from 4,805 to 7,159 tokens per second, with latency tradeoffs reported alongside the gain.
View on GitHub (opens in a new tab)Measured benchmark · Autoregressive generation
An optimization study of PyTorch inference paths. Autoregressive generation improved from 552.6 ms to 108.2 ms, while manual CUDA Graph capture reached 4.30× faster per-step execution.
View on GitHub (opens in a new tab)A GPT-2 Large scaling study that explains why four H100 nodes ran 2.4× slower than one: 93.5% of the representative step was synchronization over an NCCL socket fallback without RDMA.
View on GitHubA load-testing and telemetry project that reduced P95 latency from 98.5 seconds to 3.44 seconds, eliminated observed timeouts, and raised request success from 32% to 87%.
View on GitHubAn Airflow- and MLflow-based evaluation workflow for coding agents, with isolated execution, artifact tracking, and a tiny three-task SWE-bench Verified run used to validate the pipeline—not claim a model benchmark.
View on GitHubAn evaluation-first RAG study across 42 SEC filings and 100 FinanceBench questions, comparing retrieval, reranking, and prompt policy while improving answer correctness from 0.23 to 0.34.
View on GitHubA publishable record of a constrained fine-tuning experiment that did not establish the intended result. The repository documents the CPU-only fallback, lack of 4-bit loading, 60-step run, and limits on interpretation.
View on GitHub02 / Expertise
I connect low-level performance with the reliability of the whole system.
Benchmarking and optimization across model serving, autoregressive decoding, distributed training, quantization, and GPU execution.
PyTorch · CUDA Graphs · vLLM · H100 · NCCLControlled experiments that expose bottlenecks, quantify tradeoffs, and keep negative or inconclusive results visible.
MLflow · RAGAS · SWE-bench · Locust · RCARetrieval, agent, and inference services designed around measurable quality, typed interfaces, and reproducible behavior.
LangGraph · FastAPI · FAISS · FastMCP · PydanticRepeatable evaluation workflows with orchestration, experiment tracking, containers, tests, and CI-ready project structure.
Airflow · MLflow · Docker · GitHub Actions · pytest03 / A little about me
Before profiling models, I was investigating physical processes. The question has always been the same: what is the system actually doing?
My path runs from electronics engineering and manufacturing into data, backend software, and ML systems. I bring that hands-on perspective to building practical tools and making messy information understandable.
More on LinkedIn (opens in a new tab)The path so far
UNI — Sophisticated Electronic Assembling
Investigate production and quality problems across PCB assembly from inspection data, using root-cause analysis and process knowledge alongside production, quality, and engineering teams.
Eleos Health
Built an automated PDF-to-data pipeline, analyzed PostgreSQL datasets, and enabled A/B testing that measured a 53% improvement in documentation speed.
Freelance
Delivered Python data-collection and validation pipelines for software-development clients, producing reusable JSON and CSV datasets.
Have something in mind?
Let’s make it work.ML systems, evaluation challenges,
or the next great team.