AI · Tutorials / MenFem Daily · Episode 1 · something you keep

Inference crossover calculator

How many months until running AI on a Mac you own pays for itself against paying for API calls — at your hours, your query rate and your electricity.

I did this myself

Connor runs local models on the M4 Mini the defaults describe. The wattage is a wall reading, not a spec sheet.

65 W under load · 28 tok/s out · 500 tok/s in

Hardware and API prices are quoted in dollars. This is how they land on your bill.

4 h
116
20 /h
560
Breakeven — the month the machine has paid for itselfAPI wins
295mo
API, a month · Claude Fable
£5.62
Local, a month · 36-month write-down
£32.22

The API wins. The machine would pay for itself in month 295, past its 36-month write-down — you would replace it before it broke even.

Cumulative cost, month by month

GBP · list prices · 36-month write-down
API, cumulative Local, cumulative
Where local starts winning
Never — not at any number of hours in a day, at this query rate.
What a query costs you
£0.0023 on the API · 22s of machine time

Your numbers in the datacenter’s terms

SemiAnalysis, 14 September 2026 compared a Jetson Thor on a desk against a B300 in a rack. Once utilization is counted — a rack kept 90% busy, a device 40% — sending the work to the cloud costs 54 cents for every on-device dollar, and for a device used 12 hours a day it falls to 12%. One datacenter GPU does the work of about 7 devices. Those are their numbers. These are yours:

Your utilization
17%
API per local pound
17p
Machines for the load
1

What this counts, and what it leaves out

  • The API side is list price. No prompt caching, no batch discount, no committed-use deal. If you have one, your API month is smaller than the number here.
  • The local side is the sticker price over three years, plus electricity under load. Not counted: the model being worse, your time setting it up, a RAM upgrade, or the hours the box idles at a few watts.
  • Capacity is the number the datacenter maths hides. A rack never runs out of headroom; a Mini does. Each preset carries a tokens-a-second rate for a model that fits it, and if your demand needs more than one machine the bill says so.
  • The frame is SemiAnalysis, 14 September 2026. Their numbers are a Jetson Thor against a B300 — datacenter silicon on both sides. This page translates the shape of that argument, utilization first, onto hardware a person buys.