Skip to content
Tech News
← Back to articles

Cost of AI Model Tokens Falling Sharply as GPU Efficiency Doubles Every Two Years

read original get NVIDIA Jetson Orin Nano Developer Kit → more articles
GoKawiil Brief

An analysis argues that the cost of running machine learning models is dropping by orders of magnitude annually, driven by GPU efficiency gains that double roughly every two years—a pace not seen since early Moore's Law. The piece distinguishes proprietary models like GPT-6 Astra from open-weight models such as GLM-5.3-flash, noting that hosted and locally-run versions improve at different rates, with per-token pricing for frontier models not falling as consistently as costs for smaller models.

Why It Matters

GoKawiil's interpretation of the reporting above, not reported fact.

If token costs keep falling, the author suggests AI could shift from being a metered product to embedded infrastructure across computing within a year or two, and frontier-quality models could plausibly run locally on ordinary hardware within three to six years. This would reframe the constraint on AI adoption from raw compute cost toward model quality and access, according to the author's projection—though the piece itself frames these as likely trends rather than certainties.

Key Takeaways
Worth a Look

NVIDIA Jetson Orin Nano Developer Kit — As the article notes, frontier-quality LLMs are heading toward running locally on commodity hardware, and this compact developer kit is a great entry point for experimenting with local AI inference today. It's built specifically for running neural networks efficiently at the edge, letting you test smaller open-weight models like Qwen3 Coder on your own desk. Perfect for tinkerers wanting a hands-on feel for where local AI is headed.

See NVIDIA Jetson Orin Nano Developer Kit on Amazon → Affiliate link — we may earn a commission on purchases, at no extra cost to you. Product picked by AI based on this article; it is not a tested recommendation.

Source: jyn.dev, 2026-09-23

Published there as: “Tokens Too Cheap to Meter”

Read the original report → The summary and analysis above are GoKawiil's own, written from reporting by the source above. Facts and quotes belong to the original publisher.