Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Just 8 months in, India’s vibe coding startup Emergent claims over $100M ARR

Tesla avoids 30-day suspension in California after removing ‘Autopilot’

SpendRule Raises $2M, Comes From Stealth To Help Hospitals Track Spending

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Running AI models turns into a memory game

    18 February 2026

    Mistral AI acquires Koyeb in first acquisition to support its cloud ambitions

    18 February 2026

    Cohere launches a family of open multilingual models

    17 February 2026

    You’ve got money, you’ll travel: a16z’s hunt for the next European unicorn

    17 February 2026

    Fractal Analytics’ IPO debut signals persistent AI fears in India

    16 February 2026
  • Apps

    US court bans OpenAI from using ‘Cameo’

    18 February 2026

    These are the countries that are moving to ban social media for children

    18 February 2026

    TikTok is launching a Local streaming option in the US leveraging users’ exact location

    15 February 2026

    A Stanford student created an algorithm to help his classmates find love. Now, Date Drop is the basis of his new startup

    14 February 2026

    Airbnb plans to build AI functions for search, discovery and support

    14 February 2026
  • Crypto

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025
  • Fintech

    Cash app adds payment links so you can get paid in DMs

    11 February 2026

    MrBeast’s company buys Gen Z fintech app Step

    9 February 2026

    Stripe Alumni Raise €30M Series A for Duna, Backed by Stripe and Adyen Executives

    5 February 2026

    Fintech CEO and Forbes 30 Under 30 alum indicted for alleged fraud

    3 February 2026

    How Sequoia-backed Ethos went public while rivals lagged behind

    30 January 2026
  • Hardware

    How Recursive Intelligence Raised $335M at a $4B Valuation in 4 Months

    16 February 2026

    Nothing opens its first retail store in India

    14 February 2026

    YouTube is finally launching a dedicated app for Apple Vision Pro

    12 February 2026

    Humanoid robot startup Apptronik has now raised $935M at a $5B+ valuation

    11 February 2026

    Kindle Scribe Colorsoft is an expensive but beautiful color e-ink tablet with AI features

    6 February 2026
  • Media & Entertainment

    Apple Podcasts is getting an improved video experience this spring

    18 February 2026

    The new Amazon Fire TV interface is now rolling out in the US

    17 February 2026

    Hollywood is not happy with the new Seedance 2.0 video generator

    15 February 2026

    Designer Kate Barton teams up with IBM and Fiducia AI for a NYFW presentation

    14 February 2026

    YouTube introduces an AI playlist maker for Premium users

    14 February 2026
  • Security

    Intellexa’s Predator spyware used to hack journalist’s iPhone in Angola, investigation finds

    18 February 2026

    The European Parliament is blocking artificial intelligence on lawmakers’ devices, citing security risks

    17 February 2026

    DOJ says Trenchant boss sold holdings to Russian broker able to access ‘millions of computers and devices’

    15 February 2026

    Amazon’s Ring cancels partnership with Flock, a network of AI cameras used by ICE, feds and police

    15 February 2026

    Sex toy maker Tenga says hacker stole customer information

    14 February 2026
  • Startups

    Just 8 months in, India’s vibe coding startup Emergent claims over $100M ARR

    18 February 2026

    SpaceX Vets Raise $50M Series A for Data Center Links

    18 February 2026

    Blackstone backs Neysa in up to $1.2 billion in funding as India pushes to build domestic AI infrastructure

    16 February 2026

    As AI data centers push their power limits, Peak XV supports Indian startup C2i to fix the problem

    16 February 2026

    Twilio co-founder’s fusion power startup raises $450 million from Bessemer and Alphabet’s GV

    15 February 2026
  • Transportation

    Tesla avoids 30-day suspension in California after removing ‘Autopilot’

    18 February 2026

    Ford turns to F1 and rewards the construction of a $30,000 electric truck

    18 February 2026

    What Epstein’s files reveal about EV startups and Silicon Valley

    16 February 2026

    TechCrunch Mobility: Rivian’s savior | TechCrunch

    16 February 2026

    Aurora’s driverless trucks can now travel longer distances faster than human drivers

    14 February 2026
  • Venture

    SpendRule Raises $2M, Comes From Stealth To Help Hospitals Track Spending

    18 February 2026

    Climactic launches hybrid fund to get startups through ‘valley of death’

    18 February 2026

    African defense tech Terra Industries, founded by two Gen Zers, raises additional $22 million in one month

    16 February 2026

    How to enter a16z’s ultra-competitive Speedrun accelerator program

    16 February 2026

    India doubles state-backed venture capital, approves $1.1 billion fund

    15 February 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»Running AI models turns into a memory game
AI

Running AI models turns into a memory game

techtost.comBy techtost.com18 February 202603 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Running Ai Models Turns Into A Memory Game
Share
Facebook Twitter LinkedIn Pinterest Email

When we talk about the cost of AI infrastructure, the focus is usually on Nvidia and GPUs — but memory is an increasingly important part of the picture. As superscalers prepare to build new billion-dollar data centers, the price of DRAM chips has soared about 7 times in the last year.

At the same time, there is an increasing discipline in orchestrating all that memory to make sure the right data gets to the right agent at the right time. Companies that own it will be able to make the same queries with fewer tokens, which can be the difference between folding and staying in business.

Semiconductor Analyzer Doug O’Loughlin has an interesting look at the importance of memory chips in his Substack where he chats with Val Bercovici, Head of AI at Weka. They’re both types of semiconductors, so the focus is more on the chips than the broader architecture. The implications for AI software are also very significant.

I was particularly struck by this passage, in which Bercovici examines the increasing complexity of Anthropic direct caching documentation:

Tell it is if we go to Anthropic’s direct caching pricing page. It started as a very simple page six or seven months ago, especially as Claude Code came out — just “use caching, it’s cheaper”. Now it’s an encyclopedia of advice on exactly how much cache writes to pre-purchase. You have 5-minute levels, which are very common across the industry, or 1-hour levels — and nothing more. This is a very important element. Then, of course, you have all kinds of arbitrage opportunities around pricing for cache reads based on the number of cache writes you’ve pre-purchased.

The question here is how long Claude caches your prompt: You can pay for a 5-minute window, or pay more for an hour-long window. It’s much cheaper to pull data that’s still in cache, so if you manage it right, you can save a lot. However, there’s a catch: Each new piece of data you add to the query may display something else than the cache window.

This is complex stuff, but the bottom line is pretty simple: Memory management in AI models is going to be a huge part of AI in the future. Companies that do it well will rise to the top.

And there is much progress to be made in this new field. Back in October, I covered a startup called Tensormesh that was working on a layer in the stack known as cache optimization.

Techcrunch event

Boston, MA
|
June 23, 2026

Opportunities exist elsewhere in the stack. For example, lower down the stack, there’s the question of how data centers use the different types of memory they have. (The interview includes a nice discussion of when DRAM chips are used instead of HBM, though it’s pretty deep into the hardware.) Further up, end users figure out how to structure their model clusters to take advantage of shared cache.

As companies get better at orchestrating memory, they will use fewer tokens and inference will become cheaper. Meantime, Models become more efficient in processing each tokenpushing costs even further. As server costs come down, many applications that don’t seem viable now will start to increase their profitability.

Claude dram Exclusive game Humane inference cost memory models running turns
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleUS court bans OpenAI from using ‘Cameo’
Next Article SpendRule Raises $2M, Comes From Stealth To Help Hospitals Track Spending
bhanuprakash.cg
techtost.com
  • Website

Related Posts

Ford turns to F1 and rewards the construction of a $30,000 electric truck

18 February 2026

Climactic launches hybrid fund to get startups through ‘valley of death’

18 February 2026

Mistral AI acquires Koyeb in first acquisition to support its cloud ambitions

18 February 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Just 8 months in, India’s vibe coding startup Emergent claims over $100M ARR

18 February 2026

Tesla avoids 30-day suspension in California after removing ‘Autopilot’

18 February 2026

SpendRule Raises $2M, Comes From Stealth To Help Hospitals Track Spending

18 February 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

Cash app adds payment links so you can get paid in DMs

11 February 2026

MrBeast’s company buys Gen Z fintech app Step

9 February 2026

Stripe Alumni Raise €30M Series A for Duna, Backed by Stripe and Adyen Executives

5 February 2026
Startups

Just 8 months in, India’s vibe coding startup Emergent claims over $100M ARR

SpaceX Vets Raise $50M Series A for Data Center Links

Blackstone backs Neysa in up to $1.2 billion in funding as India pushes to build domestic AI infrastructure

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.