Close Menu
TechTost
  • AI
  • Apps
  • Crypto
  • Fintech
  • Hardware
  • Media & Entertainment
  • Security
  • Startups
  • Transportation
  • Venture
  • Recommended Essentials
What's Hot

Wikipedia blacklists Archive.today after alleged DDoS attack

Google VP warns two types of AI startups may not survive

These former Big Tech engineers are using artificial intelligence to navigate Trump’s trade mess

Facebook X (Twitter) Instagram
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer
Facebook X (Twitter) Instagram
TechTost
Subscribe Now
  • AI

    Sam Altman would like to remind you that people use a lot of energy too

    22 February 2026

    ‘Toy Story 5’ takes aim at creepy AI toys: ‘I’m always listening’

    21 February 2026

    Great news for xAI: Grok is now very good at answering questions about Baldur’s Gate

    21 February 2026

    UAE’s G42 partners with Cerebra to deploy 8 exaflops of computers in India

    20 February 2026

    Why these startup CEOs don’t think AI will replace human roles

    20 February 2026
  • Apps

    Apple’s iOS 26.4 arrives in public beta with AI music playlists, video podcasts and more

    22 February 2026

    India’s Sarvam launches Indus AI chat app as competition heats up

    21 February 2026

    Remember HQ? “Quiz Daddy” Scott Rogowsky is back with TextSavvy, a daily mobile game show

    21 February 2026

    As the browser war heats up, Chrome is adding new productivity features

    20 February 2026

    Google says its AI systems helped prevent Play Store malware in 2025

    20 February 2026
  • Crypto

    Hackers stole over $2.7 billion in crypto in 2025, data shows

    23 December 2025

    New report examines how David Sachs may benefit from Trump administration role

    1 December 2025

    Why Benchmark Made a Rare Crypto Bet on Trading App Fomo, with $17M Series A

    6 November 2025

    Solana co-founder Anatoly Yakovenko is a big fan of agentic coding

    30 October 2025

    MoviePass opens Mogul fantasy league game to the public

    29 October 2025
  • Fintech

    InScope raises $14.5M to solve financial reporting pain

    20 February 2026

    OpenAI deepens India push with Pine Labs fintech partnership

    19 February 2026

    Cash app adds payment links so you can get paid in DMs

    11 February 2026

    MrBeast’s company buys Gen Z fintech app Step

    9 February 2026

    Stripe Alumni Raise €30M Series A for Duna, Backed by Stripe and Adyen Executives

    5 February 2026
  • Hardware

    Joseph C Belden: Last Chance for Innovators to Earn Scaling Privileges

    20 February 2026

    At a critical time, Snap is losing a top spec executive

    20 February 2026

    Freeform Raises $67M Series B to Scale Laser AI Production

    19 February 2026

    India’s Sarvam wants to bring its AI models to phones, cars and smart glasses

    19 February 2026

    Google debuts $499 Pixel 10a

    18 February 2026
  • Media & Entertainment

    Google adds music-making capabilities to its Gemini app

    21 February 2026

    Disrupt 2026 Super Early Bird pricing expires in 1 week

    20 February 2026

    YouTube’s latest experiment brings its AI chat tool to TVs

    20 February 2026

    OpenAI, Reliance partner to add AI search to JioHotstar

    19 February 2026

    SeatGeek and Spotify are teaming up to offer concert ticket discounts within the music platform

    19 February 2026
  • Security

    Wikipedia blacklists Archive.today after alleged DDoS attack

    22 February 2026

    Error on student admissions website exposed children’s personal details

    21 February 2026

    Ukrainian man jailed for identity theft that helped North Koreans get jobs at US companies

    21 February 2026

    Cellebrite cut off Serbia citing misuse of its phone unlocking tools. Why not others?

    20 February 2026

    FBI says ATM ‘jackpot’ attacks on the rise, hackers net millions in stolen cash

    20 February 2026
  • Startups

    Google VP warns two types of AI startups may not survive

    22 February 2026

    Co-founders behind Reface and Prisma join hands to improve on-device model inference with Mirai

    21 February 2026

    Nominations for the Startup Battlefield 200 are now open

    21 February 2026

    The OpenAI mafia: 18 startups founded by graduates

    20 February 2026

    Nvidia deepens early-stage push into India’s AI startup ecosystem

    20 February 2026
  • Transportation

    These former Big Tech engineers are using artificial intelligence to navigate Trump’s trade mess

    22 February 2026

    Rivian owners will soon be able to access vehicle controls using their Apple Watch

    21 February 2026

    Lucid Motors is cutting 12% of its workforce as it pursues profitability

    21 February 2026

    New York puts the brakes on robotaxi expansion plan

    20 February 2026

    AI data center boom fuels Redwood’s energy storage business

    20 February 2026
  • Venture

    Ali Partovi’s Neo appears to upgrade the throttle model in low dilution terms

    21 February 2026

    Peak XV Raises $1.3B, Doubles In AI As Global India VC Competition Heats Up

    21 February 2026

    General Catalyst commits $5 billion to India over five years

    20 February 2026

    Reload wants to give your AI agents a shared memory

    20 February 2026

    This VC’s best advice for building a founding team

    19 February 2026
  • Recommended Essentials
TechTost
You are at:Home»AI»Silicon Valley bets big in ‘environments’ to train agents AI
AI

Silicon Valley bets big in ‘environments’ to train agents AI

techtost.comBy techtost.com22 September 202509 Mins Read
Share Facebook Twitter Pinterest LinkedIn Tumblr Email
Silicon Valley Bets Big In 'environments' To Train Agents Ai
Share
Facebook Twitter LinkedIn Pinterest Email

For years, Big Tech CEOs have inaugurated AI agents who can use autonomous software applications to complete people. But take today’s AI Agents for a rotation, be it the Openai Chatgpt agent or the Perplexity comet and you will quickly realize how limited the technology is. Making AI agents can get a new set of techniques that the industry is still discovering.

One of these techniques carefully simulates the workplaces where agents can be trained in multiple-step duties-known as reinforcement environments (RL). Similarly in the way the data sets supplied the last wave of AI, the RL environments begin to look like a critical element in the development of factors.

Researchers, founders and investors AI tell TechCrunch that top AI laboratories now require more RL environments and there is no lack of newly formed businesses that hope to supply them.

“All the big AI laboratories build RL environments at home,” said Jennifer Li, a collaborator at Andreessen Horowitz, in an interview with TechCrunch. “But as you can imagine, creating these sets of data is very complicated. So AI laboratories also consider third -party suppliers who can create high quality environments and ratings.

The push for the RL environments has crying a new category of well -intentioned newly established businesses, such as engineering and primary intellect, aiming to drive the space. Meanwhile, large data labeling companies such as Mercor and Surge say they are investing more in RL environments to keep up with industry shifts from static data sets in interactive simulations. Big workshops are also thinking of investing in large $ 1 billion in RL environments Next year.

The hope for investors and founders is that one of these newly established companies emerges as a “AI scale for environments”, referring to the $ 29 billion Powerhouse data marking.

The question is whether the RL environments will really push the borders of AI progress.

TechCrunch event

Francisco
|
27-29 October 2025

What is a RL environment?

At their core, the RL environments are educational reasons that simulate what an AI agent would do in a real software application. A founder described their construction recent interview “Like the creation of a very boring video game.”

For example, an environment could simulate a Chrome browser and work an AI agent with the purchase of a pair of socks on Amazon. The agent is scored by his performance and sent a reward signal when he succeeds (in this case, buying a worthy pair of socks).

While such work sounds relatively simple, there are many places where an AI agent could escape. Navigation in the developing menus of the website may be lost or buy too many socks. And because developers cannot predict exactly what a mistake will turn an agent, the environment itself must be durable enough to capture any unexpected behavior and deliver useful comments. This makes the construction environments much more complex than a static data set.

Some environments are quite complex, allowing AI agents to use tools, internet access, or use various software applications to complete a given task. Others are closer, with the aim of helping an agent learn specific tasks in Enterprise software applications.

While RL environments are the hot thing in Silicon Valley at the moment, there is a lot precedent for using this technique. One of Openai’s first projects in 2016 was the construction ”Gyms rl“Which were quite similar to the modern perception of the environments. The same year, Google Deepmind’s Alpha The AI ​​system struck a world champion in the board game, Go. He also used RL techniques in a simulated environment.

What is unique to today’s environments is that researchers are trying to create AI agents using computers with large transformer models. Unlike Alphago, which was a specialized AI system that works in a closed environment, today’s AI agents are trained to have more general opportunities. AI researchers today have a stronger starting point, but also a complex goal where more can go wrong.

A full of field

AI data labeling companies such as Scale AI, Surge and Mercor are trying to meet the moment and create RL environments. These companies have more resources than many newly established businesses in the field, as well as deep relationships with AI Labs.

Surge Edwin Chen CEO tells TechCrunch that he has recently seen a “significant increase” in demand for RL environments within AI laboratories. Surge – which he created reportedly Revenue of $ 1.2 billion Last year from collaboration with AI Labs such as Openai, Google, Anthropic and Meta – recently turned a new internal organization specially tasked with building RL environmental, he said.

The closure behind the Surge is Mercor, a startup of $ 10 billion, which has also worked with Openai, Meta and Anthropic. Mercor puts investors for RL Business Building environments for specific tasks, such as coding, healthcare and law, according to the marketing material observed by TechCrunch.

Mercor CEO Brendan Foody told TechCrunch in an interview that “few understand how big the opportunity around the RL environments is.”

The AI ​​scale has used to dominate the data label, but has lost ground since Meta invested $ 14 billion and hired its CEO. Since then, Google and Openai have fallen on the AI ​​scale as a data provider and even the start is facing competition for work with data labeling in the Meta. But still, the scale is trying to meet the moment and build environments.

‘This is just the nature of the business [Scale AI] It is means, “said Chetan Rane, the scale of AI’s product for agents and RL environments.” The scale has proven its ability to adapt quickly. We did this in the early days of autonomous vehicles, our first business unit. When Chatgpt came out, the AI ​​scale adapted to it. And now, once again, we are adapting to new border venues such as agents and environments. ”

Some younger players focus exclusively on environments from the beginning. Among them is engineering, a starting start about six months ago with the bold target of “automation of all jobs”. However, co -founder Matthew Barnett tells Techcrunch that his business starts with RL environments for AI encoding agents.

Mechanize aims to provide AI laboratories with a small number of powerful RL environments, Barnett says, instead of larger data companies that create a wide range of simple RL surrounding. At this point, boot offers software engineers $ 500,000 For the construction of an environment of RL – much higher than an hourly contractor could earn work on a AI or Surge scale.

Mechanize has already worked with humanity in RL environments, two sources familiar with the issue told TechCrunch. Mechanize and Anthropic refused to comment on the partnership.

Other newly established companies bet that RL environments will have an influence outside AI laboratories. Prime Intellect – a boot supported by researcher AI Andrej Karpathy, Founders Fund and Menlo Ventures – aims at smaller RL environments.

Last month, Prime Intellect started a Rl hub surroundings, aimed to be a “hugged person for RL surroundings.” The idea is to give open source developers to access the same resources that the large AI laboratories have and sell these developers access to computing resources in the process.

Training generally capable factors in RL environments can be more computing than previous AI training techniques, according to Prime Intellect Will Brown. Along with the newly established companies that create RL environments, there is another opportunity for GPU providers that can supply the process.

“The RL environments will be too big to dominate any company,” Brown said in an interview. “Part of what we do is just try to build good open source infrastructure around it.

Will it score?

The open question around the RL environments is whether the technique will escalate like previous AI training methods.

Aid learning has powered some of the biggest jumps in AI in the past year, including models such as Openai’s O1 and OPENAI’s Claude Opus 4, are particularly important discoveries, because the methods previously used to improve AI models now show reduced release.

The environments are part of Ai Labs’ largest stake in RL, which many believe will continue to lead to progress as they add more data and computational resources to the process. Some of the Openai researchers behind O1 told TechCrunch that the company initially invested in AI reasoning models that were created through RL and Compute Time-Time-because they thought it would be fine.

The best way for the RL scale remains unclear, but the environments look like a promising candidate. Instead of simply rewarding chatbots for text answers, they let agents work in simulations with tools and computers available. This is much more intense, but possibly more rewarding.

Some are skeptical that all these RL environments will get rid of. Ross Taylor, a former AI researcher with Meta who co -founder of general reasoning, tells Techcrunch that RL environments are prone to rewarding hacking. This is a process in which AI models cheat to get a reward, without really doing the work.

“I think people underestimate how difficult it is to escalate the environments,” Taylor said. ‘Even the best available to the public [RL environments] They usually do not work without serious modification. ”

The head of OpenAi engineering for API business, Sherwin Wu, told a recent podcast That was “short” in the newly established RL Environmental Businesses. Wu noted that it is a very competitive space, but also that the AI ​​research is evolving so quickly that it is difficult to serve AI’s laboratories well.

Karpathy, a primary intellect investor called RL environments a possible discovery, has also expressed attention to the RL area wider. To one Post in xHe raised concerns about how much the progress of AI can be squeezed by RL.

“I am swollen in environments and techniques of interactions, but I am a Bearish in enhancing learning in particular,” Karpathy said.

UPDATE: A previous version of this article refers to mechanical work as mechanical work. Has been informed to reflect the official name of the company.

agent agents bets big environments Human learning open Research Rl Scale ai Silicon train Valley
Share. Facebook Twitter Pinterest LinkedIn Tumblr Email
Previous ArticleTrump says Lachlan and Rupert Murdoch could invest in Tiktok deal
Next Article Fueled by India’s small businesses, UK Fintech Tide becomes a Unicorn supported by TPG
bhanuprakash.cg
techtost.com
  • Website

Related Posts

These former Big Tech engineers are using artificial intelligence to navigate Trump’s trade mess

22 February 2026

Sam Altman would like to remind you that people use a lot of energy too

22 February 2026

‘Toy Story 5’ takes aim at creepy AI toys: ‘I’m always listening’

21 February 2026
Add A Comment

Leave A Reply Cancel Reply

Don't Miss

Wikipedia blacklists Archive.today after alleged DDoS attack

22 February 2026

Google VP warns two types of AI startups may not survive

22 February 2026

These former Big Tech engineers are using artificial intelligence to navigate Trump’s trade mess

22 February 2026
Stay In Touch
  • Facebook
  • YouTube
  • TikTok
  • WhatsApp
  • Twitter
  • Instagram
Fintech

InScope raises $14.5M to solve financial reporting pain

20 February 2026

OpenAI deepens India push with Pine Labs fintech partnership

19 February 2026

Cash app adds payment links so you can get paid in DMs

11 February 2026
Startups

Google VP warns two types of AI startups may not survive

Co-founders behind Reface and Prisma join hands to improve on-device model inference with Mirai

Nominations for the Startup Battlefield 200 are now open

© 2026 TechTost. All Rights Reserved
  • About Us
  • Contact Us
  • Privacy Policy
  • Terms and Conditions
  • Disclaimer

Type above and press Enter to search. Press Esc to cancel.