Kill your own product when the experiment beats it
Most founders treat the thing they built as the thing they have to make work. Every new signal gets read as evidence for the plan.

Like this episode?
Get real founder strategies for the AI era. Delivered weekly.
Free weekly newsletter · No spam
Eugene Cheah spent two years building his own AI model, then ran a weekend experiment on the side. Over that launch weekend the experiment made more revenue than the main platform, so he killed the original product and rebuilt the company around it. Featherless AI now serves more than forty thousand open source AI models and reached multiple seven figures in ARR within about a year.
The hardest part was not technical. As he puts it, the realization was that people wanted these models more than they wanted his model, and that he had been holding his own mission back.
Eugene Cheah's team had built something genuinely novel: an open source foundation model under the Linux Foundation, trained across more than 200 languages, with an architecture that made AI inference dramatically cheaper. Their seven billion parameter model beat Llama's equivalent. The problem was that almost nobody was asking for it.
Along the way they solved a constraint of their own making. Customers were fine-tuning thousands of models on their platform, the industry norm was one GPU per model, and they could not afford a thousand GPUs. So they built a system that swaps models on and off GPUs on demand, bringing a cold model online in about five seconds. Most inference providers keep a fixed list of under a hundred models standing by, because loading one can take thirty minutes on hardware costing eighty dollars an hour.
Then someone on the team asked whether the same technology would work for Llama and Mistral. They shipped it as an experiment. Over the launch weekend it earned more than the platform they had spent two years on. Eugene renamed the company and went all in.
In this interview, Eugene explains why he priced a flat monthly rate while the rest of the AI industry charged per token, how stripping the technical explanation off the homepage kept improving conversion until they removed their own research from the top of the page, and why competing for the long tail of open source AI models beats fighting a hundred providers over the top hundred.
Featherless AI founder Eugene Cheah killed two years of work on his own open source AI model after a side experiment made more money than the main platform over a single launch weekend, then grew Featherless to more than forty thousand open source AI models and multiple seven figures in ARR within about a year by serving the long tail of models no other AI inference provider would host.
Most founders treat the thing they built as the thing they have to make work. Every new signal gets read as evidence for the plan.
Usage-based pricing looks fair. You charge for what people consume, costs scale with revenue, and nobody overpays.
Most validation advice tells you to go ask people. Run interviews, describe the idea, see if they say yes.
Go where the demand is. Serve the popular use case, the biggest segment, the models and categories everyone already searches for.
Why did Eugene Cheah abandon two years of work on his own AI model after a weekend experiment made more money than his main platform?
The experiment made more revenue over its launch weekend than the platform he had built for two years. He concluded that people wanted these models more than they wanted his model, and that holding on was holding his own mission back.
What is GPU hot-swapping and why does it matter for AI inference?
Standard tooling can take 30 minutes to load an AI model onto a GPU, so providers keep a small fixed list standing by on hardware costing around eighty dollars an hour. Featherless swaps models in on demand in about five seconds, so one GPU can serve many models.
How many open source AI models does Featherless AI host compared to other inference providers?
More than forty thousand, against under a hundred for a typical provider. Eugene's goal is all three million models on Hugging Face.
Why did Eugene Cheah choose flat-rate pricing when the AI industry charges per token?
Their earlier product had too many prices and confused buyers, and teams could not tell a CFO what AI would cost. A fixed monthly rate removed bill shock and made procurement possible. Major labs later shipped fixed-price coding plans.
How did removing information from the Featherless AI homepage improve conversion?
The original page explained the underlying research and the speculative decoding technique. Each time they de-emphasized the explanation, conversion improved, until they removed even their own models from the top of the page. Most customers already knew which model they wanted.
How did Featherless AI get its first customers?
Through Reddit communities like LocalLlama and Ollama and through Discord, where fine-tuners were already asking how to run models they could not host themselves. Word of mouth, partnerships, and a Hugging Face integration followed.
Why does Featherless AI compete for the long tail instead of the most popular models?
The top hundred models attract ten or more providers each. Beyond that list Featherless is typically the only provider. Eugene would rather own a quarter of a market nobody contests than fight a hundred competitors for the top slice.
What business advantage does hosting unpopular AI models create?
Models with fewer than a million requests a month are not worth a dedicated GPU for most providers, so they refuse them. Hot-swapping makes them economic, which turns a cost problem into an exclusive inventory position.
What does Eugene Cheah mean about leaving half the world out of AI?
His grandmother speaks seven languages, none of them English or Chinese, and AI is dominated by both. More than half the world speaks neither. Featherless hosts first language models for Uganda, Thailand, Indonesia, and the Philippines.
Eugene Cheah [00:00:00]:
As much as we wish that we were the number one model in open source space, the reality is we were not. As an experiment, we created a platform that is now known as Featherless AI and over the weekend and launched it, we had more revenue than the main platform. The reality is people want these models more than your model.
Eugene Cheah [00:00:21]:
And it was a realization that I was holding my own mission back.
Omer Khan [00:00:26]:
Hey, welcome to the SaaS podcast. I'm Omer Khan and this is the show where I sit down with real founders and dig into how they actually built their SaaS companies. I've had almost 500 of these conversations now and I put out a new one every week to help you build and grow your startup. If that sounds useful, hit subscribe or check out SaaS Club I.O.
Omer Khan [00:00:45]:
To learn more. My guest today is Eugene Cheah, founder of Featherless AI, an AI startup that began as a weekend experiment. Within about a year, that experiment grew into a platform serving more than 40,000 open source AI models and generating multiple seven figures in ARR. In this interview, Eugene breaks down why he killed two years of work to go all in on that experiment, why he bet on flat rate pricing when the rest of the AI industry charged per token, how giving users less information on their website actually improved conversions and and how Reddit and Discord drove his first million in ARR.
Omer Khan [00:01:26]:
So I hope you enjoy it. All right, Eugene, welcome to the show.
Eugene Cheah [00:01:29]:
Glad to be here.
Omer Khan [00:01:30]:
So tell us about Featherless. What does the product do? Who's it for?
Eugene Cheah [00:01:35]:
So Featherless AI is trying to build the world's largest platform where you have instant access to all the world's collection of open models. Today we have 40,000 models and we are trying to scale it to all 3 million models that you can see online are hugging face. We believe in a future where anyone can have access to any AI model without another person telling them what they can or cannot have access to.
Eugene Cheah [00:02:00]:
We want to make it easy to just play around with all these AI models that will be building the future that we see ahead of us.
Omer Khan [00:02:07]:
Sweet. Give us a sense of the size of the business, where are you in terms of revenue, team, customers, that stuff.
Eugene Cheah [00:02:14]:
So we just closed our Series A previously which is 20 million series A with Airbus Ventures and AMD Ventures with the goal of scaling to 10 million ARR by the end of the year. And the scheme size is around 30 people geo distributed across the entire world. So we have teams members all the way in Singapore, Toronto, Europe, get to close on Japan.
Eugene Cheah [00:02:39]:
But once I have that, I literally have the most extreme Case of West coast and east coast all around the world.
Omer Khan [00:02:46]:
Yeah. So tell me, when did you found the company and when did you hit the first million in ARR?
Eugene Cheah [00:02:52]:
We founded the company in around two and a half years, closer to two years, but it was doing a very different thing. So for those who knew about history more is that fabulous As a company started under a previous name, Recurso, because most of the members that we had, we were building on the RWKV open source project.
Eugene Cheah [00:03:19]:
So this was a Next Generation Foundation AI model that has properties that make inference cost 1000x cheaper to do inference and had the potential to completely change the way we use AI. But the issue that we had at the technology was that it was unproven at larger scales. So the team was iterating and building on this and our hallmark at that point in Time was a 7 billion parameter model that beat llama 7B.
Eugene Cheah [00:03:51]:
And we created a platform where you can fine tune our models with your own data to push the frontier of low cost inference. Essentially what ended up happening was that we had a platform where people were fine tuning thousands of models and the norm in AI was one GPU, one model and we couldn't afford 1000 GPUs at that time.
Eugene Cheah [00:04:22]:
And so we developed a new inference platform where you can access any of the RWKB model and swap between them on the fly to serve all these new open source RDMKV models. Then someone in the team asked the technology for this part to be able to serve thousands of models is not exclusive to rwm. Kiwi, right? Could we get this to work for LLAMA and Mistral, which was the popular models at the time?
Eugene Cheah [00:04:50]:
As much as we wish that we were the number one model in open source space, the the reality is we were not. And so as an experiment we created a platform that is now known as Federless AI. To say okay, because since our platform is so laser focused on rwkv, Ricasso and all these fine tuning systems, why don't we just create a new platform as an experiment to just run any LLAMA administrator model, one click access under a new system that is lightweight serverless inference, hence featherless AI.
Eugene Cheah [00:05:32]:
And using this technology that we have developed and launch it. And that was approximately one year ago and over the weekend as we launched it, we had more revenue than the main platform. And so that was the point where we pivoted into Faedalus. Because instead of the original mission of the RWKD group, which was to make AI accessible to everyone regardless of Language or compute.
Eugene Cheah [00:06:04]:
We realized that through Faedalus we can maintain the same mission. It's just sadly not exclusive to our model, but to make all the AI models accessible, regardless of language and computer.
Omer Khan [00:06:15]:
Right? Okay, so this started out as an experiment just to see can we use this technology with a different model. And very quickly you realize there's more demand and more traction here than with what we've been doing for the last year or so or whatever. Explain a little bit about the GPU hot swapping. A lot of people listening might not understand what the big deal is.
Omer Khan [00:06:44]:
Just explain you did the one to one, right, One model, one GPU and you didn't have the money, so you were swapping. But practically how easy or hard is it to do that? How much of a big deal was it that you were able to do something like this and hot swap?
Eugene Cheah [00:06:59]:
So the thing is, today, if you look at the landscape, right, despite the fact that we have 3 million AI models on hugging Face, for example, the average inference provider only provides you less than 100 models. And that's because the existing software we have in place can take 30 minutes just to load an AI model and start it up and have everything running.
Eugene Cheah [00:07:25]:
30 Minutes is for the biggest model. Some of the smaller models may be 10 minutes, give and take. And all this means that for a provider, especially with some of these GPU hardware being literate servers, like $80 an hour, worth millions of dollars, you do not want to waste time getting them to switch between models. And also for customers, when you make an API request, you're not going to say, you're not going to accept, hey, please come back 30 minutes later.
Eugene Cheah [00:08:02]:
To give an analogy as to why this is important, what end up happening is that providers end up having a fixed list of models. And the experience is more like going to a bakery. It's already all pre baked, it's already all pre standby. This standby takes $80 plus an hour. And when you want that piece of bread that you see, you ask for it, you get it on the spot.
Eugene Cheah [00:08:31]:
You cannot ask for anything else. Instant need on the spot. And that was the issue that a lot of inference providers do. They only can limit their inventory to keep the flow moving. But to be able to hot swap, change this equation, it brings it closer to, let's say Starbucks, where you can get your Frappuccino with a piece of cream and custom seasoning on top.
Eugene Cheah [00:09:00]:
That is possible because we can deploy your AI model when you make that request within the acceptable time frame for us, it was five seconds. If, let's say, the model isn't online in our system, when you make the request within five seconds, it comes online and you're chatting with it. And that changed fundamental unit economics in a lot of ways.
Eugene Cheah [00:09:23]:
We don't need to stand by as much GPUs because we can swap things in and out dynamically. And also a GPU can be serving customer A with model A for one hour and it can switch to customer D with model D in the next second when customer A finishes. So it allows us to serve more models in the same hardware as anyone else.
Eugene Cheah [00:09:50]:
And that was basically what we built as a solution to the problem that we had where we had thousands of.
Omer Khan [00:09:57]:
Fine tuned models and not enough money.
Eugene Cheah [00:10:00]:
Exactly. And framing it as we knew the market would be hungry for that actually gave it more credit than what was happening. Because we were just trying to experiment on how to get our line of RBKV models, which is an open source model with the analytics foundation, to work. It has a lot of interesting properties and we still do research on this, but the challenge was to get customers to use it.
Eugene Cheah [00:10:37]:
And that's why we create fine tuning. And that's why we created this problem. And then we end up solving the problem. And then we weren't even sure or certain whether there was demand in the market for it. But that's why we did an experiment.
Omer Khan [00:10:53]:
What was the reason for the experiment? Just explain that. I think you said you tested with LLAMA initially. What were you hoping to learn from this experiment?
Eugene Cheah [00:11:06]:
The question is, is there enough people who are interested or have demand for this? And what we found was that there were smaller communities within the local LLM committee, for example, or localama and Olama committee, who were constantly experimenting and fine tuning with different AI models for different use cases and even in some cases for different languages for those who are outside of us to support their own region.
Eugene Cheah [00:11:42]:
And they were constantly having issues of trying to get these models up and running. Because even if I can afford my $10,000 GPU, my friend cannot run the model because it's on a 10,000 GPU that's maybe inside my house. And so the unresolved question was that would there be enough demand? But we saw the messages, we saw people asking, hey, on Reddit as finetuners created this model, how do I run this model then?
Eugene Cheah [00:12:21]:
How can I have access to it? If only there was a way for me to have instant access to it. When we realized we built that technology, even though it's for Arabic theory and we saw the demand that there was for Llama and Mistral. That was where we decided to just do the experiment. So it was less intentional and more like we saw there was enough demand in Reddit and it was like there was no harm.
Omer Khan [00:12:51]:
Right. So it was more about, let's just get it out there, let's see if there's interest in doing this. If there's another application for this technology that we've built, I think it's worth also explaining. Just like the idea of an inference provider, it's probably a term that a lot of people listening to the show might not be familiar with.
Omer Khan [00:13:11]:
And so all of these models, these open source models are available, right? You can go to hugging face and access them. The problem becomes, as you said exactly, is where the hell am I going to actually run this? And that's where an inference provider comes in.
Eugene Cheah [00:13:27]:
Correct? An inference provider basically, essentially is a provider online that takes these models and provides access for you at a certain price. That's the simplest way to frame it. There are multiple major providers that do so for major open source models. There are also providers that do so for closed source models. So by stretching the definition, for example, if let's say you're using OpenAI models or entropic models, you are actually using them as an inference provider, but they also happen to be the creator of the model.
Eugene Cheah [00:14:11]:
Another example for OpenAI, Entropic, for example, you could go to Google Cloud, AWS and Azure to run anthropics model. And OpenAI's model, in this case Google Azure or AWS is the inference provider. Well, OpenAI and Anthropic is the model creator. That's for closed source. For open source that's where you have the Facebook llama model, which was the initial wave of popular open source model Mistral.
Eugene Cheah [00:14:52]:
But more recently there's a much larger variety. We are talking about, let's say quant, Deepseek, that most people may have heard of Zai's JRM 5.2 model. And the options have kind of grew over time. Instead of having a handful of main models right now, we are getting new open source models that are pushing the frontier literally on a monthly basis.
Eugene Cheah [00:15:26]:
And it's almost like clockwork. Every month there'll be a new major open source model. And people are constantly trying that as replacements for the existing close source model. And there was a gap when we first started. We knew that as open source model progresses that the gaps between closed source and open source will close. And hence why we decided to build this inference Platform to support the open source models instead of the closed source models.
Omer Khan [00:16:02]:
Yeah. So let's talk about that experiment. You kind of put it out there, suddenly started seeing that there was a big appetite in the market for a solution like this. Walk me through how you then decided that this was a thing that you were going to focus on. You didn't just wake up the next day and say let's host all the millions of models on hugging face and go after that.
Omer Khan [00:16:33]:
Which is probably the vision now. But when you were running that experiment and you saw this, what was the next step or what was the direction you started thinking about?
Eugene Cheah [00:16:45]:
I would say it's more actualization of the mission. The mission that we had was always to make AI accessible. One of the stories that I give an example I give is that my grandma speaks seven languages, she doesn't speak English or Chinese. And if you look at the existing AI landscape, it's English or Chinese dominated. And in Southeast Asia there is hundreds of languages.
Eugene Cheah [00:17:22]:
In fact, in both South America and Africa, even India alone has hundreds of languages in those spaces. For more than 50% of the world, they don't speak English or Chinese. To me, when I founded the company, my mission is to make AI accessible regardless of language or compute. Because to me what was scary was that if AI has all this potential and promise and it's doing all these wonders that will change the world economy and our future economy will be built on AI.
Eugene Cheah [00:18:07]:
What horrifies me as well is that we may leave half the world out of it and I'm not affected. I speak English, I'm perfectly fine, but it's my relatives and that half of the world that I see that I do not want to live out. And so to me, that was the mission, the core of the mission to make AI accessible.
Eugene Cheah [00:18:31]:
And that's why I went into the RWK RE project, because we were one of the first AI models under the Linux Foundation. We were one of the first models that trained over 200 languages that is trying to make AI more accessible through languages. We were also the first to try to change the attention mechanism to be something that is much cheaper to do inference.
Eugene Cheah [00:18:54]:
So for an AI model that you can run on your laptop, we were at that point in time the frontier of running AI models in a non English or Chinese language. And because of that I was so focused on how do we improve our model to get there. And hence the whole company was built around that. But the reality that I had to accept was that the demand was not for these smaller models.
Eugene Cheah [00:19:33]:
That could run, let's say on your mobile phone. The demand was for the smarter ChatGPT like models that people are more used to on ChatGPT itself and because the demand is there respectively it was more of like how do we make this more accessible? And it was federalist. Subsequently when it started moving, we realized it came as a reflection.
Eugene Cheah [00:20:04]:
It's more like the reality is people want these models more than your model. And it was a realization that I was holding my own mission back. And since then we went off on a clearer path. It's not just about rwkt, it's about all open source model. Since then, for example, we were one of the first and few providers for Uganda's first language model or Thai's first major language model or the Indonesia's first fine tuned language model.
Eugene Cheah [00:20:42]:
We are also working with a team that is fine tuning specifically for Philippines, their target model. And so we still maintain our mission. But it was no longer RWKB, it was you could be fine taming Lama, Mr. Quan, it doesn't matter. We will try to make it accessible on our platform. And that was to me the biggest change.
Eugene Cheah [00:21:07]:
Well, technically if you rewind back, the mission is the same. It's just we were trapped in our thinking.
Omer Khan [00:21:13]:
Tell me about the pricing decision that you made. We're in a world where anybody who's using these types of models is going to be used to paying some kind of usage based model. That's what anthropic and OpenAI are pushing everyone towards. And you decided actually we're going to try to flat rate and we're going to give you access to pretty much everything.
Omer Khan [00:21:43]:
What was the reasoning behind that?
Eugene Cheah [00:21:44]:
I think there's a few reasonings behind it. For one, when we did the previous iteration based on rbkv, we had models of different sizes and we had too many models of different sizes I would argue. And what ended up happening is people was like why is there so many prices? And they were very confused on on what size to use and for what prices.
Eugene Cheah [00:22:18]:
And that is still true today. There was one. So we had a lot of confusion there. We also saw that for a lot of non AI native teams they were struggling to tell their boss that said hey, I want to introduce AI to the company and then their CFO will say will be okay, how much will it cost?
Eugene Cheah [00:22:44]:
Maybe $10, maybe $1,000, maybe 10,000, I don't know until we try it. And the average response is like what? And it slowed down procurement and realizing all those constraints. And also at the Same time reading the horror stories of like, hey, I thought I was only going to spend $5 each, charge $500. We realized that there is a place for a fixed rate provider that will guide you from bill shocks, that, that will design the product around this fixed rate provisioning.
Eugene Cheah [00:23:26]:
And at that point of time, we realized there was no one in the market that did that. And so that was part of the experiment because at the start we were going to launch with 5,000 models. And I'm like, I don't want to set up a pricing table for 5,000 models. I'm just going to set a flat rate.
Eugene Cheah [00:23:46]:
And behind the scenes we adjusted things to make sense. Bigger models run slower, smaller model runs faster. People kind of get it. And, and that kind of worked because we were positioning this for experiments, for personal needs. We were not positioning the subscription models plan at $25 a month or $10 a month for hey, run your entire production system of thousands of requests in parallel.
Eugene Cheah [00:24:18]:
Here it was more like run your individual chat and request or experiments. And in hindsight we were right. Because if you look fast forward today, almost all the major labs now have a fixed pricing plan for coders, for developers. They call this the coding plan. $200 A month, you can make as many requests up to a certain rate limit, but you'll never get the bill.
Eugene Cheah [00:24:47]:
Shock. And that allowed a lot of developers to use AI models within companies that would say, hey, I cannot have this token based pricing. That was something that when at that time we were the first to do it, but now it became slightly more and more common. And if you look at other SaaS products like monitoring databases, you also realize that there is a place for both usage based pricing and subscription by a specific capacity based pricing.
Eugene Cheah [00:25:25]:
And they exist side by side. So it was just a case of reading AI. No one was doing it, we just did it and that's it.
Omer Khan [00:25:32]:
Okay, so from what I understand, you got traction pretty quickly with that experiment. How long did it take from the time you did that first experiment to hit the first million in ARR?
Eugene Cheah [00:25:47]:
If you talk about from the company's launch, more than one and a half years.
Omer Khan [00:25:51]:
Yeah, from the experiments launch from Federation,.
Eugene Cheah [00:25:55]:
It was pretty much we were already on track within the first few months.
Omer Khan [00:26:00]:
Okay.
Eugene Cheah [00:26:00]:
And so it was a case of like this was so obvious and we just went along with it.
Omer Khan [00:26:07]:
And then when you closed on 2025, you guys were already beyond a million in ARR at that point. Right. You already had enough serious traction to know that this was where you were Going to focus your energy?
Eugene Cheah [00:26:20]:
Yes.
Omer Khan [00:26:21]:
Okay, so I'm interested. I think the pricing issue is really valid. You don't want to have 4000 models available and you're trying to create a pricing list of each one. And there's just. It's a headache for you. It's complicating for anybody trying to figure out which model they should pick and why and how much is it going to cost and so on.
Omer Khan [00:26:42]:
So you solve that problem with simple pricing. But when you and I talked previously, you had said that one of the other issues you had was also the messaging was too complicated about what Featherless did, and that actually hurt your conversions. So tell me about that. What was the problem that you were struggling with there?
Eugene Cheah [00:27:05]:
So that was back to the origin story, right? Arabic. So we still, to be clear, we still do research on it, but on FedEx platform itself, right? What we did was that when we first launched, we're like instant access to 4,000 models now. Today, right now it's 14,000 models. And the original page was like powered by RWKV.
Eugene Cheah [00:27:32]:
Because people like to ask like, hey, how do you manage to get inference to this cheap and this affordable? Like, how does this make sense? And that's where we shared that inside the description and details. How is this possible? We run to your models with a speculative decoder, which is an inference technique where a small model drafts an answer and the big model accepts or reject it.
Eugene Cheah [00:27:56]:
And this doesn't affect your quality, it's just a technique. Today, every inference provider use it, including OpenAI and Google. Google was the one who wrote the original paper, but we were doing it differently with other Kitty models. So we explained it. This is how we lower your inference cost by 10x and so on and so on. And we realized over time that people were just spending too much time asking those questions.
Eugene Cheah [00:28:27]:
And in a lot of cases, as more and more inference provider was growing, it was just a case of like, hey, this is the model, this is the price, so be it, you can run it. And so we started doing experiments where we de emphasize how we did everything. We removed like the top header on optimize with this, yada, yada yada, we put it down, it improved conversion, we further put it down and part by part, literally remove all the explanation.
Eugene Cheah [00:28:57]:
This is a surprise. These are the models that we have and it improved conversion. And then it reached to a point where right now if you go to feathers at the top page, it's like you don't even see the very models that we create anymore. We used to put that at the top and it was like, nope, it does reduce conversion.
Eugene Cheah [00:29:14]:
And so that's basically realizing that most people don't care actually. They just know what models they want. They heard about it and they want to try. So let's say when Deep SEQ happened, they don't care what was the architecture changes. They just said, deep seq, I want to try it. So for example, for us, our research into linear attention, which we are very vested in, for example, when Quin 3.5 came out to us, we were extremely happy that because they used parts of our research to lower their model inference cost by 75%.
Eugene Cheah [00:29:58]:
So our research part of it end up going into the quant architecture. No one asked about it other than a few really dedicated researchers and dedicated AI folks that were really passionate about it. But for the 90% of the customers that were running the model, they had no idea. And so we're just like, you want model A, you get model A.
Eugene Cheah [00:30:27]:
That's it. This is the price.
Omer Khan [00:30:29]:
I mean, in many ways it reminds me of that story of Steve Jobs when the ipod came out and every other provider at the time was Talking about these MP3 players and how much storage they had and blah, blah, blah, and all the technical specs of these things. And then he came along and he was like, well, you know the saying about you can have so many songs in your pocket.
Omer Khan [00:30:55]:
That's what really people cared about. They didn't care about how the thing was built or, you know, the tech, Tech specs. People. People might want to kind of, you know, out of curiosity, look at that. But that's not the main thing. And you know, when I go to featherless AI website now, I see one API key, instant access.
Omer Khan [00:31:14]:
And initially when I saw that for the first time, I was like, oh, so this is kind of like an open router type thing, right? I have one API key and you let me access like OpenAI and Anthropic and all of this stuff. And it's not. It's a very different type of offering now that I understand it better.
Omer Khan [00:31:34]:
But at the end of the day, does it really matter? Right? It's like people just want to be able to use a model as easily as possible. And if you make that, you know, you make that promise and you give, it's a few clicks away and the deployment is a few seconds away. That's probably the thing they care the most about.
Omer Khan [00:31:52]:
Yeah.
Eugene Cheah [00:31:53]:
The amount of times that folks think that we are an open router, I'm like, no, we are not an open router. We actually host these models. Open router routes models to providers like us. It's the other way around. But like you said to the customer, it's like they get what they want and that's it.
Omer Khan [00:32:14]:
So this vision of taking all of these open source models on hugging Face and hosting them and being the inference provider, I mean that's like a bold goal, right? I mean you have like 40,000 models today, but how many models are there on hugging face today? Right? 2, 3 Million? Right, so that's a lot. Like how do you like, why not focus on just like the most important ones?
Omer Khan [00:32:57]:
Does this go back to your vision about making it more accessible for everybody? Or is there something else behind you wanting to kind of say we want to host millions of these models?
Eugene Cheah [00:33:11]:
It's both because like the accessibility piece is an important statement because especially for models that will not receive more than a million requests a month, most providers will not bother hosting. Just doesn't justify if they need to stand by a GPU 24. 7 For it. Because our technology allows that differentiator, we can do that. But beyond that, it was also like an understanding that the direction moving forward.
Eugene Cheah [00:33:48]:
Because when I see AI and let's say when I was young, I imagined the Jetsons where basically every family have one robot that's helping it very personalized to them. Not one AI model or five AI models by a handful of companies that consolidates all the AI traffic and decides what can run or not run or what is supported.
Eugene Cheah [00:34:22]:
And so even when we were building for our own line of model, I was very clear saying I do not want another billionaire to decide who can use what model. We wanted to let people fine tune and iterate.
Omer Khan [00:34:35]:
From what I understand, you kind of basically just look at any signals. The signal you look on hugging face is if there's a model that's even had what like 100 downloads, you auto onboard that and you make that available.
Eugene Cheah [00:34:48]:
We right now increase the threshold to around 1000, but users can request.
Omer Khan [00:34:52]:
Okay, okay, okay.
Eugene Cheah [00:34:54]:
And actually as we go down that path. We go down that path. Is that what we also realize is that there is a demand for AI to be more personalizable, more personalized to individual companies. Back to the Jetson's example. And well, it's not the case today. The signals are really there. For example, we work with, let's say various folks companies and we see it happening now like for example att, Shopify, they fine tune their own models for their own company use case att supported like their customer support and proprietary telco switching system.
Eugene Cheah [00:35:39]:
Shopify supported their own Shopify configuration language. And we see this trend, more and more companies are creating their own model. And this really just speaks to the indicator that if you're only hosting 100 models, you're not enough for this future. There'll be thousands and millions of it. And that's how we tie the vision to the commercials.
Omer Khan [00:36:06]:
How did you get the word out? I know initially when you did the experiment, you were reaching out to people and making them aware on places like Reddit and Discord that you had this new solution available. But beyond that, where did the growth come from? How did you acquire customers?
Eugene Cheah [00:36:24]:
During the original launch, it was really through all those Discord, Reddit and so on to acquire the initial phase and then after that it was through word of mouth and then subsequently through partnerships and channels. So for example, we are integrated directly on Hugging Face and we are not the only inference provider to do so. And back to the point of is there an intentional differentiator?
Eugene Cheah [00:36:52]:
It's like us providing more models than anyone else allows us to differentiate from the crop of inference model. If you go to the number one model, there are like 10 providers. We are just one in 10. But if you go to any other model beyond the top 100 list, for the models that we host, typically we are the only provider.
Eugene Cheah [00:37:18]:
And so for customers that go there and say, hey, I want to use this model, let's say Matt Jammer, which is a medical fine tune, or Cisco's security model, which Cisco created for security use cases and you go to Hal Inface. The only provider was Fantasy. It's through these partnerships that also people find through Discover Earth and it's through these integrations as well.
Eugene Cheah [00:37:44]:
So it's like we started from, just like you said, the Reddit and Discord community, but we went beyond that into its partnerships and integrations and through essentially word of mouth and events.
Omer Khan [00:37:56]:
Right, so basically you're kind of basically trying to capture the long tail of the models on Hugging Face, where maybe the big players won't necessarily care about, but given the way that the technology that you have and the way that you can hold these different models, host these different models and do the GPU hot swapping, it's a pretty achievable thing that you can deliver to people.
Eugene Cheah [00:38:24]:
Yes, because I mean, I think this is what baffles me that not enough people are paying attention to this in terms of my competitors. But the thing is, I even get talks on it and I do have some places like conferences where I disclose some of the numbers. If you look at the pie chart per se, the top hundred is like the top 60 percentile of the models.
Eugene Cheah [00:38:58]:
All the obvious names are that deepseed, coin, etc. But our bottom 25% are slices that are so small that I can't even see the individual colors in the pie chart. And if the inference market by estimates are like it's trillions of dollars worth, would you rather compete with 100 other providers at the top 50% or would you rather compete with no one at the 25% slice?
Eugene Cheah [00:39:31]:
And from a business standpoint it actually made sense because you'll be like, I can get, I potentially can grow into a 25% of the 3 trillion dollar market that no one else competing with. I'm down.
Omer Khan [00:39:44]:
So you raised a series A recently. That was what was like 20 million or so?
Eugene Cheah [00:39:49]:
Yes.
Omer Khan [00:39:49]:
And then, so what's the mission now? How are you going to spend that money?
Eugene Cheah [00:39:54]:
Scale scaling up more access to the models. We are now at 40,000 models. We are still a long distance away from all 3 million and we intend to just make them all instantly accessible.
Omer Khan [00:40:08]:
And then I used to rely on hugging face as the main sort of distribution.
Eugene Cheah [00:40:11]:
Not so much. It's more of like how some people get to know us. We realized that initially most people are fine tuning models and putting them on hugging face. One of the growing segments for us.
Omer Khan [00:40:28]:
Is.
Eugene Cheah [00:40:31]:
Actually hosting fine tuned models that companies are fine tuning themselves. Another major growing segment these days is really, really more like we've been organizing more events and conferences and in particular within Europe. And a lot of companies are starting to pay a lot of attention to running open source model because they are not guaranteed access to the closed source model.
Eugene Cheah [00:40:58]:
As with what happened recently with the anthropic fable, people cut off access to the best model. And the funny thing is when you cut off the best closed source model, the version behind it is almost equal to the best open source model. So for these companies it's like what's the difference here? And when they came to that realization that hey, not only is the open source model 10 times cheaper, but about the same in terms of what we get, more and more companies started switching.
Eugene Cheah [00:41:31]:
So that has made a huge push recently as well.
Omer Khan [00:41:33]:
Yeah, it's a very, very interesting space. Anyway, we should wrap up. Let's get onto the lightning round. I've got five quick fire questions for you. Ready? What's one of the best pieces of business advice you've received?
Eugene Cheah [00:41:46]:
Start with why once you know why A lot of things become clearer. Why you're doing this in the business.
Omer Khan [00:41:52]:
What book would you recommend to our audience and why?
Eugene Cheah [00:41:55]:
The same books that we buy.
Omer Khan [00:41:58]:
Simon Sinek. How did I know you were going to say that? What's the best money you've ever spent on your business?
Eugene Cheah [00:42:06]:
It was actually before the business. What got us here in part on the RWKB project and the whole long chain is that I basically could afford to set aside that $10,000 worth of money to buy GPUs and to do all these experiments before ChatGPT, and that kind of let us stick here. Was it a waste of money?
Eugene Cheah [00:42:36]:
Yeah, it was kind of a scary spend.
Omer Khan [00:42:38]:
Well, funny how it turned out. What's your favorite personal productivity tool or habit?
Eugene Cheah [00:42:44]:
This sounds dumb, but I end up overusing it more often than I like WhatsApp or Telegram to send myself a message because I'm hopping across so many devices that it's just useful for it to work. I know there's a better solution. And I know people are like doing AI agents, and I have my AI agents in the WhatsApp channel and Telegram, but having that to just jot a quick memo and notes still happens to be like 60% of all my notes.
Omer Khan [00:43:19]:
And finally, what's one of your most important passions outside of your work?
Eugene Cheah [00:43:23]:
Outside of my world, it will be physics and space in general, or anything aeronautical as well. This is what when I was young I wanted to be a pilot and then when I realized there wasn't really much opportunity and I basically want to go into physics and in space and there wasn't really much opportunity to be an astronaut, especially in Southeast Asia.
Eugene Cheah [00:43:52]:
So I was like. And so that's how I end up. Not related to computers. That's the short version.
Omer Khan [00:44:00]:
Love it. Well, thank you so much for joining me. It's been a pleasure. If people want to check out featherless, they can go to featherless AI. And if folks want to get in touch with you, what's the best way for them to do that?
Eugene Cheah [00:44:11]:
On my Twitter picocreator, I do have a substack as well. Tech Talk cto. Otherwise, just drop me an email as well. It's not a healthy guest. Eugene at Federalist.
Omer Khan [00:44:25]:
Awesome, Eugene, thank you so much for joining me. Been a pleasure. Wish you all the best.
Eugene Cheah [00:44:29]:
Thank you for having me here.
Omer Khan [00:44:30]:
My pleasure. Cheers.

Farzad Rashidi, Respona
Farzad Rashidi is the co-founder of Respona, a company that helps brands get cited in AI answers across ChatGPT, Perplexity, and Google AI Overviews. He first came on the show back in episode 323, when Respona was a self-serve outreach tool doing a few hundred thousand in ARR. Then the classic bootstrapped trap set in. Churn caught up with new business, and every customer they won was offset by one they lost. For years Farzad tried to fix it the way most founders do, by adding more features to make the product stickier. Nothing moved. The real reason customers left was not missing features. It was that they never had the time to do the work the tool required. The turning point came in early 2025. A marketing agency CEO haggled over an $800-a-month license, then offered to pay per result instead. That one conversation nudged Farzad toward a service-as-software model: do the work for the customer, charge for the outcome, and use the software in the back end. That first customer now spends around $65K to $70K a month, and in twelve months the company 4x'd the revenue it had spent six years building. What makes this a service-as-software story rather than a slide back into agency work is what came next. Farzad demoted the self-serve SaaS on the homepage, productized the service into fixed tiers with no negotiation, and rebuilt a software layer (a client portal, a publisher network, and a brain in the middle) on top of the manual delivery. We also dig into the actual playbook for getting a brand cited in AI answers, from finding lookalike publishers to building a surround-sound presence around the models. I hope you enjoy the conversation.

Marius Meiners, Peec AI
Marius Meiners is the co-founder and CEO of Peec AI, a platform that helps marketing teams track how their brands appear on AI search tools like ChatGPT, Perplexity, and Gemini. After studying economics, working in venture capital and M&A at PwC, and joining Antler's Berlin cohort, Marius found himself with no team, no idea, and four years removed from writing any code. Then in late October 2024, ChatGPT launched search. Marius saw it and decided this was going to change everything. The smartest SEO experts in the world were already obsessed with it. The signal was loud. So he turned to AI search optimization as the wedge - a category that would explode as marketers scrambled to figure out how to get cited by AI assistants. He vibe coded the first prototype with V0 in a day and a half. Eight customers signed letters of intent based on it. Antler wrote a 100K check. His CTO joined and built the real product in six weeks. Peec launched in February 2025. Then came the bet. Their biggest competitor had raised five times more money and was chasing the world's biggest brands. Marius made the opposite call. Peec priced at 85 euros while competitors charged over 500. For six months, Marius and the team ate two-euro canned food every day, wondering if the mid-market AI search optimization play would ever pay off. Today Peec has over 2,000 customers, $8.6 million in ARR, and a team of 55. All in 14 months. AI search optimization went from speculation to a live revenue channel - 20% of Peec's own conversions now come through AI search itself.

Girish Redekar, Sprinto
Girish Redekar is the co-founder and CEO of Sprinto, an autonomous compliance platform that helps companies prove they're handling data securely. Before Sprinto, Girish and his co-founder spent two to three years trying to build their first startup. They had no programming background, so they taught themselves to code at 28 because they couldn't afford to hire developers. The first few ideas went nowhere. A job search engine. A resume matching tool. None of them got traction. Then they built RecruiterBox, a simple CRM for hiring. It launched before Stripe even existed, so their payment system was absurd. Customers had to click a PayPal link, swipe their card, and get credits that depleted daily. When the credits ran out, they'd go back to PayPal and swipe again. It was terrible. But customers kept doing it anyway. That was the clearest signal of finding product-market fit Girish had ever seen - not from analytics or feedback forms, but from watching people jump through hoops to keep paying. They bootstrapped RecruiterBox to over 2,500 customers and single-digit millions in ARR. Then they sold it. Not because it was failing, but because it got too comfortable. They felt like they were the bottleneck, and the business would grow faster in someone else's hands. The second time around, Girish took a completely different approach to finding product-market fit. He'd experienced the pain of SOC 2 compliance firsthand at RecruiterBox, spending months and tens of thousands of dollars on consultants. So when he started Sprinto, he made a rule - no code until the idea was validated. He used The Mom Test framework, ran 15-20 customer interviews, and then paid auditors to audit his non-existent company. Ten times. Each audit, he built a little more product behind the scenes. By the tenth, he knew exactly what to build and had confirmed the consulting service could actually become software. For go-to-market, Girish tried 20 different channels. Seventeen failed. The three that worked were founder communities (Slack groups, WhatsApp groups), VC portfolio programs with startup discounts, and Google (paid + SEO). His framework - harvest existing demand rather than create new demand - helped him focus on places where people were already looking for solutions. Now AI is changing the game from three directions at once. It's changing Sprinto's product, it's changing how customers operate internally, and it's creating new security threats from the outside. The problems that didn't exist two years ago are now driving a whole new wave of demand. Today, Sprinto generates eight-figure ARR with over 3,000 customers across 75 countries. The team has grown to 350 people and they've raised $32 million.