Tech News
← Home  ·  All topics

Ai Inference

9 GoKawiil briefs on this topic

Epoch AI: Cost of matching AI performance fell 47% per quarter since 2023

A new Epoch AI report by Emberson and Roodman finds that the price of achieving a given level of AI performance has dropped roughly 47% per quarter over the past three years, a 13-fold annual decline. The report cites OpenAI's o3, which cost about $0.30 per question to score 75% on GPQA Diamond in January 2025, compared with a much newer model reaching similar scores for a fraction of a cent by mid-2026.

Crusoe raises $3.9B at $30.9B valuation to build modular AI data centers

Crusoe, the AI infrastructure firm behind Oracle's giant Abilene, Texas campus, has closed a $3.9 billion funding round led by Atreides Management, Valor Equity Partners and Mubadala, pushing its valuation to roughly $30.9 billion. The company is using the capital to expand into smaller, factory-built facilities called Spark, designed for AI inference rather than massive training clusters.

AI industry shifts focus from training to inference workloads in 2026

Major AI labs have moved their attention from building ever-larger models to running inference—the process of using trained models to generate text, code, and images. This shift is driven by growing real-world use of large language models, the rise of reasoning models that repeatedly reprompt themselves, and autonomous AI agents that run continuously rather than just responding to single queries. Amazon Web Services, for instance, has split inference tasks between its Trainium chips and Cerebras's wafer-scale hardware.

Qualcomm shares jump 4% after AWS data center chip partnership, $4B warrant deal

Qualcomm's stock rose 4% on Tuesday following news that it will supply customized silicon for Amazon Web Services' AI infrastructure, with a focus on inference workloads. As part of the deal, Qualcomm granted Amazon warrants to purchase 25 million shares at $161.26 each, valuing the arrangement at roughly $4 billion.

Gimlet Labs raises $300M Series B led by Andreessen Horowitz for AI inference optimization

Gimlet Labs has closed a $300 million Series B funding round led by Andreessen Horowitz, with participation from Sapphire Ventures, Menlo Ventures, Arm, Samsung Ventures, Tiger Global and others. The company says it has added billions in contracted revenue and gigawatts of datacenter pipeline since March, and is scaling toward hundreds of megawatts of managed capacity, just five months after its Series A.

Analysts urge enterprises to rearchitect memory and storage for AI inference

Tirias Research founder Jim McGregor argues that AI is not one workload but millions of varied ones, meaning data centers can no longer treat memory and storage as secondary hardware. He says inference and agentic AI demand purpose-built systems capable of continuous, real-time data ingestion, caching, and movement rather than legacy infrastructure designed for training workloads.

Artificial Analysis and Liquid AI launch benchmark for phone-based small AI models

Artificial Analysis has begun benchmarking small AI models that run directly on mobile phones, focusing on models that fit within 8 GB of memory after quantization, including KV cache at 8K context. The effort combines intelligence benchmarks with real on-device inference data gathered in partnership with Liquid AI, whose measurement methodology Artificial Analysis says it has independently validated.

Z.ai unmasks mystery model Ox Alpha as open-weight GLM-5.3-Flash

An unidentified free model called Ox Alpha appeared on OpenRouter and quickly drew massive usage from developers who couldn't determine its origin, sparking days of speculation about which lab built it. Z.ai revealed the model was actually its own GLM-5.3-Flash, running entirely on Chinese chips, and priced at 15 cents per million input tokens and 50 cents per million output tokens, with a launch discount cutting that in half through September 9.

IBM previews mainframe chip that natively runs both Z and Arm instructions

At Hot Chips 2026, IBM unveiled an early look at a future mainframe processor featuring 11 IBM Z cores on a 2-nanometer process, clocked above 5.7 GHz, capable of natively executing both z/Architecture and AArch64 instructions on the same cores without emulation. The chip will include specialized accelerators for AI fraud detection, I/O, compression, cryptography and sorting, plus a large cache hierarchy topping out near 3.5 GB of virtual L4 cache. IBM hasn't named the chip or its host system but expects it around 2028, likely powering a future z18-class mainframe.