Computation log field-note

Efficient AI Systems

A 12-week research and engineering program investigating small language models on CPU-only infrastructure, from raw transformer inference to a production Go service on Kubernetes.

How much useful AI can we build under strict compute constraints?

Efficient AI Systems is a research monorepo documenting a structured, experiment-driven investigation into running, measuring, evaluating, and deploying small language models (SLMs) on CPU-only infrastructure — from raw transformer inference up through a production-style Go service running on Kubernetes with full observability.

The program runs across four phases and twelve weeks: understanding what actually happens when an SLM runs on a CPU, measuring the tradeoffs that emerge when models are compressed and compared, building a reliable and observable service around one, and synthesizing all of it into an answer to the real question — under which technical, economic, and operational constraints does a smaller, more efficient model become the better engineering choice.

github.com/alessandrobessi/efficient-ai-lab