All included infrastructure company Infinity announced a $15 million raise at a $100 million valuation on Monday from investors including Touring Capital, Principal VC and researchers from companies such as OpenAI and Anthropic.
The startup makes software to make it easier for artificial intelligence chips to run AI models. A big reason why Nvidia has become the top player is not only its high-performance chips, but also its CUDA (Compute Unified Device Architecture) software, which allows its GPUs (originally designed to run graphics) to act as general-purpose CPUs. The largest AI development frameworks PyTorch and TensorFlow are built on top of CUDA. This allows developers to write their apps in popular languages like Python, use these main AI frameworks, and their apps will, by default, run on Nvidia chips.
Most of these application-level startups wouldn’t have the resources or expertise to write their own cores—the low-level software that operates chips—and port their applications to other AI chips. So, Infinity is trying to create an alternative CUDA kernel software that works with any type of chip, including SRAM, GPU, phone chips, and shrink arrays. Infinity is part of a new wave of startups trying, product by product, to chip away at Nvidia’s market dominance.
Infinity attempts to create a universal inference library that will run on all chips, allowing those chips to automate the replication of state-of-the-art research results.
Infinity was launched last year by Jeremy Nixon, a former researcher at Google Brain and creator of the AGI House hacker network community. Nixon told TechCrunch that he decided to start this company because he was obsessed with the idea of ”automated invention” — the belief that “AI systems can actually be a meta-technology.” He had invented a machine learning algorithm called Omega, he said, which essentially generated new machine learning algorithms and automatically evaluated them in a feedback loop.
That success got him thinking about other cases where this approach might work, and he turned to hardware, believing that automated systems could also create the low-level code, such as kernels and so on, needed to help the chips run more efficiently.
Infinity’s Ignition AI research agent is meant to write the low-level code needed for AI inference on alternative Nvidia chips. It tests, debugs and measures how fast the hardware performs with the code and automatically rewrites the code if needed to improve performance. The system is self-optimizing, meaning it is constantly learning and improving. It also adapts to different chip architectures, regardless of proprietary designs, Nixon says. The result is what Infinity claims is a CUDA-level software stack.
Customers include AI chip maker (and would-be Nvidia challenger) D-Matrix, and Infinity is in talks with other major chip and cloud companies, Nixon said.
However, humans are in the loop, providing high-level direction while the agent does more of the tedious grunt work. In a case studythe startup found that the agent works much faster than a human alone, reducing what could have been years or months to hours or days. Infinity does not charge upfront license fees. Instead, a reduction in performance gains and cost savings is required by measuring token changes per second.
Infinity currently has 26 employees, including those in design, operations and engineering.
When you purchase through links in our articles, we may earn a small commission. This does not affect our editorial independence.
