
Eugene Cheah, Featherless AI
The Weekend Experiment That Killed His Own AI Model
Eugene Cheah's team had built something genuinely novel: an open source foundation model under the Linux Foundation, trained across more than 200 languages, with an architecture that made AI inference dramatically cheaper. Their seven billion parameter model beat Llama's equivalent. The problem was that almost nobody was asking for it. Along the way they solved a constraint of their own making. Customers were fine-tuning thousands of models on their platform, the industry norm was one GPU per model, and they could not afford a thousand GPUs. So they built a system that swaps models on and off GPUs on demand, bringing a cold model online in about five seconds. Most inference providers keep a fixed list of under a hundred models standing by, because loading one can take thirty minutes on hardware costing eighty dollars an hour. Then someone on the team asked whether the same technology would work for Llama and Mistral. They shipped it as an experiment. Over the launch weekend it earned more than the platform they had spent two years on. Eugene renamed the company and went all in. In this interview, Eugene explains why he priced a flat monthly rate while the rest of the AI industry charged per token, how stripping the technical explanation off the homepage kept improving conversion until they removed their own research from the top of the page, and why competing for the long tail of open source AI models beats fighting a hundred providers over the top hundred.























