I build and evaluate language-model features at Microsoft — Copilot across Word, Excel, and PowerPoint — and I lead evaluation for the Office Copilots. Most of my work comes down to one unglamorous question: how do you know whether the thing you shipped is actually good? Benchmarks, assertions, judge models, and the failure taxonomies that come out the other side.
Before Microsoft I spent four years as a data scientist at Penske Logistics shipping freight-pricing and driver-safety models, and three years before that writing C++ for automotive infotainment. Eleven years of putting models in front of people who will notice when they are wrong.
writing
I’m building a transformer from nothing — no PyTorch, no NumPy, just Python lists and loops — and then adding back everything real inference systems need: batching, ragged batching, a KV cache, sampling. Each post ends by proving the new version still matches the old one.
papers
Office Comprehension Benchmark (OCB): A Benchmark for Document Comprehension across Office Documents
preprint · arXiv forthcoming · first author
elsewhere
linkedin · github · scholar · orcid · firoz@firozshaik.com