# Before you buy another Mac, measure the whole request.

An evaluation protocol for MLX, Ollama, LM Studio, and exo that keeps useful output, waiting time, context, and memory in the same decision.

Published: 2026-10-09 | Updated: 2026-10-09
Author: /init editorial
Method: AI-assisted, source-checked
Canonical: https://init.news/notes/apple-silicon-inference-evaluation/

## The answer

Compare local inference with a fixed workload and a versioned configuration, then separate loading, prompt processing, generation, and end-to-end completion. Test cold and warm requests, realistic context, and expected concurrency. Choose the runtime or hardware that meets your task's quality and latency requirements; a tokens-per-second number alone cannot establish that fit.

## Write the task before the benchmark

A short chat response, a long-document extraction, and a tool-using agent place different demands on the same machine. Start with a handful of representative inputs and observable success criteria: valid structured output, supported answers, correct tool arguments, or a completed transformation. Preserve failed and timed-out requests in the results. Excluding them can make a faster configuration look more useful than it is.

For a controlled runtime comparison, hold the base model and workload steady where formats permit. Record conversion and quantization differences explicitly; an MLX artifact and a GGUF artifact with similar labels are not automatically identical. For a product decision, compare each stack's best acceptable configuration, but label that as a whole-system comparison rather than an isolated runtime result.

Run condition | Keep fixed or record | Decision it informs
--- | --- | ---
Cold request | Model residency, load policy, artifact, prompt, and output limit. | Whether occasional use spends too long waiting for the first useful response.
Warm request | Cache state, repeated prefix, context length, and decode settings. | Whether the everyday interactive path meets its response budget.
Long context | Actual token counts, prompt-processing time, peak memory, and quality. | Whether extra context is usable for the intended task.
Concurrent requests | Arrival pattern, request count, queue time, and per-request failures. | Whether one service can support expected users or parallel agents.
Sustained batch | Run order, duration, power mode, memory pressure, and background load. | Whether the system remains useful beyond a flattering single request.

## Read the measurements your runtime actually exposes

Ollama documents separate response fields for model loading, prompt evaluation, and output evaluation, with timing in nanoseconds. Streaming responses carry usage in the final chunk. Save that terminal record, and measure client-observed first-token and completion time separately. Server generation duration does not include every source of waiting experienced by your application. [1]

LM Studio's v0 API documents time-to-first-token, generation speed, model information, and runtime metadata; the same page recommends v1 for new projects. Choose the API version deliberately and capture the fields it actually returns. Do not silently compare different definitions of latency across endpoints or discard missing statistics as zero. [2]

MLX LM documents prompt caching, a configurable prompt-processing step size, and a rotating key-value cache. These settings can change memory use and behavior, so include them in the configuration record. Test cache reuse as its own condition rather than letting one candidate receive a warmed prefix and the other start empty. [3]

## A compact record for every attempt

- Identity: test-case ID, attempt ID, timestamp, machine and operating-system version, runtime version, exact model artifact, and quantization.
- Inputs: prompt hash, template version, input-token count, context limit, output limit, sampling settings, and any enabled reasoning or tool protocol.
- State: loaded or unloaded model, cache policy, concurrency, memory pressure, power mode, and relevant background work.
- Results: client-observed first-token time, completion time, server timings with units, output-token count, peak memory if available, stop reason, and error class.
- Usefulness: task score, evidence or expected-output comparison, parser success, tool-call validity where relevant, and a retained output for inspection.

## Add a second machine only for a demonstrated constraint

exo's current repository describes automatic device discovery and distributed execution, with separate setup and RDMA requirements. Its documentation is the place to verify the actual machines, network, operating system, and interconnect supported by the chosen release. [4] Treat a multi-device test as a distinct system: record the topology and compare useful throughput, per-request waiting, and recovery when a worker disappears.

Run repeated trials, retain the individual observations, and report the sample size before choosing a summary statistic. Set the acceptable quality and waiting-time thresholds before seeing the winner. No local hardware was benchmarked for this guide; it supplies a protocol for deciding whether to tune the existing stack, change the runtime, or add capacity.

## Sources

- [Ollama: API usage metrics](https://docs.ollama.com/api/usage) — [1] Load, prompt, generation, cached-token, and streaming usage fields and nanosecond timing units.; checked 2026-10-09
- [LM Studio: REST API v0](https://lmstudio.ai/docs/developer/rest/endpoints) — [2] Documented timing and runtime metadata; explicit recommendation to use v1 for new projects.; checked 2026-10-09
- [MLX LM: README](https://github.com/ml-explore/mlx-lm) — [3] Prompt caching, rotating KV cache, configurable prefill step size, and configuration tradeoffs.; checked 2026-10-09
- [exo: official repository and setup guide](https://github.com/exo-explore/exo) — [4] Current distributed-inference architecture, automatic device discovery, installation and RDMA prerequisites; no performance claims adopted.; checked 2026-10-09
