I build and evaluate language-model features at Microsoft — Copilot across Word, Excel, and PowerPoint — and I lead evaluation for the Office Copilots. Most of my work comes down to one unglamorous question: how do you know whether the thing you shipped is actually good? Benchmarks, assertions, judge models, and the failure taxonomies that come out the other side.

Before Microsoft I spent four years as a data scientist at Penske Logistics shipping freight-pricing and driver-safety models, and three years before that writing C++ for automotive infotainment. Eleven years of putting models in front of people who will notice when they are wrong.

writing

I’m building a transformer from nothing — no PyTorch, no NumPy, just Python lists and loops — and then adding back everything real inference systems need: batching, ragged batching, a KV cache, sampling. Each post ends by proving the new version still matches the old one.

No matching items

all posts →

papers

Office Comprehension Benchmark (OCB): A Benchmark for Document Comprehension across Office Documents

preprint · arXiv forthcoming · first author

PPT-EVAL: A Benchmark for Computer-Use Agents on PowerPoint Tasks

ICML 2026 · code released

all papers →

elsewhere

linkedin · github · scholar · orcid · firoz@firozshaik.com