Now
Member of Technical Staff
Working on the engineering problems behind serving large language models efficiently and reliably.
Engineer, researcher, occasional writer
I work on LLM inference, multi-agent learning, and the systems that make powerful models practical.
I’m a computer scientist who likes figuring out how things work under the hood. Right now, I work on LLM inference at Cohere, where I spend most of my time thinking about memory, latency, throughput, and all the systems details between a model and the words appearing on your screen. I also do multi-agent reinforcement learning research with Zhijing Jin at the Vector Institute, studying what happens when you put intelligent agents in a world together and ask them to cooperate, compete, and communicate. I build things, run experiments, and occasionally write about whatever I find interesting.
Current interests
These days, most of my work is around LLM inference at Cohere, where I think about memory, latency, throughput, and the systems behind serving large models efficiently.
I also do multi-agent reinforcement learning research at the Vector Institute with Zhijing Jin, focused on building environments for studying how agents communicate, coordinate, compete, and learn together.
KV caches, memory pressure, latency, throughput, and the serving layer behind fast models.
Studying how agents coordinate, adapt, communicate, and learn in shared environments.
Building tools and infrastructure that make ML experiments easier to run, inspect, and scale.
Experience
Now
Working on the engineering problems behind serving large language models efficiently and reliably.
Research
Currently researching multi-agent reinforcement learning and coordination in learned systems.
2025
Worked on pretraining and finetuning generative recommender models for merchant-facing recommendations.
2025
Previously worked on biological foundation models and large-scale experiments for DNA sequence modeling.
Earlier
Built NLP classification systems, fall-detection pipelines, 3D deep learning models, simulation assets, and an LLM-driven avatar system.
Selected work
Text worlds for interactive agents
At its core, Word Play is about making multi-agent environments easier to build and understand: worlds where agents can move, communicate, cooperate, compete, and reason from natural-language observations.
Featured project
Built an avatar system using NVIDIA Omniverse, Audio2Face, Python, PyTorch, Docker, gRPC, Hugging Face, and AWS to support installation-related customer service workflows.
Writing
A look at KV-cache compression and how to recover exactness while keeping the efficiency wins.
Read article
Exploring how neural networks can be interpreted through decision tree structure.
Read article
A step-by-step guide to the pieces behind a small language model implementation.
Read articleOther projects
Newsletter