Why per-token AI pricing breaks at sensor-data volume
Per-token pricing is a good deal for occasional questions. A plant is not an occasional user: machines stream data every second and cameras inspect every part.
Every AI workload has a break-even point. Below it, paying per token is cheaper than running your own capacity. Above it, owned capacity wins, and the gap grows with every request. For office work, many teams never cross that line. In manufacturing, much of the work crosses it early.
Why plant workloads are different
Three things make manufacturing volumes unusual:
- They are continuous. IoT sensors and PLCs produce data around the clock, across every shift and weekend.
- They scale with the plant, not with headcount. Connecting one more line adds hundreds of signals and several cameras, without adding a single user.
- They are repetitive. The same checks run thousands of times a day: is this vibration pattern normal, is this weld acceptable, which step is this operator on.
Per-token pricing turns each of those checks into a line item. Finance sees a bill that rises with every connected machine, and projects stall at the pilot stage because nobody can forecast the cost of scaling them.
What owned capacity changes
With models running on GPUs in your own cloud or at the plant edge, the cost is the hardware and its operation, not the number of questions. Several things keep that cost down:
- Small models first. Most routine checks never need a large model.
- Capacity that follows shifts. GPU capacity scales with production patterns and overnight batch analysis rather than sitting idle.
- Shared hardware. Several models and tasks share the same processors.
- Edge processing. Video is analysed where it is recorded, so you are not paying to move terabytes around.
The more lines you connect, the cheaper each answer becomes, instead of the other way round.
Where managed AI still makes sense
Low-volume, public-only work, such as summarising supplier news or public regulation, is often cheaper per token. A hybrid setup keeps that work on managed models under a monthly cap, and moves any workload to owned capacity once it crosses its break-even point.
Cost is not the only reason
At very low volumes, pay-per-use can win on price. But trade secrets, export-controlled designs and camera footage of your people often justify private AI before the numbers do. The good news for manufacturers is that, at plant volumes, privacy and cost usually point the same way.
To see where your own workloads sit, try the break-even calculator on XePlatform.