For a decade, the software-as-a-service playbook was written in the language of infinite scalability. The dream was simple: build a feature once and sell it a thousand times with near-zero marginal cost. In this world, adding a new user was essentially free for the provider, allowing traditional SaaS firms to enjoy gross margins that hovered comfortably between 80% and 90%. But for the new wave of AI-first companies, that dream is colliding with the reality of the token. Every time a user prompts a model, a server hums, a GPU burns electricity, and a bill from a cloud provider grows. The marginal cost of a user is no longer zero; it is a variable expense that scales in lockstep with usage.

The Token Tax and the Death of the Seat

The financial architecture of AI-first enterprises looks fundamentally different from the legacy SaaS era. Current data indicates that AI-first companies are operating with gross margins in the 50% to 60% range. This significant dip is the direct result of the token economy. In a traditional model, a seat-based license provided a predictable revenue stream regardless of whether the user logged in once a month or a thousand times a day. In the AI era, a power user is not a high-value customer; they are a cost center. When usage increases, the cost of goods sold increases immediately, squeezing the margin.

This economic pressure is forcing a rapid migration away from traditional pricing. The percentage of AI companies utilizing seat-based pricing has dropped from 21% to 15% in just one year. The industry is realizing that charging per head is a dangerous game when the underlying cost of the service is tied to compute. To survive, companies are moving toward mixed models that balance stability with scalability.

Take Candor, an AI-powered interview platform, as a primary example of this shift. Candor has rejected the all-you-can-eat seat model in favor of a hybrid approach. Their Standard plan requires a base monthly fee of 75 dollars, which grants the user 5 projects. However, to protect their margins from high-volume users, they charge an additional 7 dollars per interview. By decoupling the platform access fee from the actual execution of the AI task, Candor ensures that as the customer derives more value from the tool, the company is compensated for the increased compute cost.

This shift is not without its risks, particularly during the customer acquisition phase. In the traditional SaaS world, a free trial cost the company almost nothing. For an AI-first company, a free trial is a direct liability. Candor reports that a single free trial provided without a credit card costs the company approximately 68 dollars in actual compute expenses. If that trial includes the first additional project, the direct cost to acquire and activate a single lead jumps to roughly 77 dollars before a single cent of revenue is generated. This transforms the free trial from a low-risk marketing tool into a high-stakes financial investment, requiring a much more precise calculation of customer lifetime value versus acquisition cost.

The Value Ceiling and the Retention Paradox

If the cost of compute creates a hard floor for pricing, the perceived value to the customer creates the ceiling. The challenge for AI founders is finding the sweet spot between these two markers. Pricing cannot drop below the cost of the tokens, but it cannot exceed the value the customer believes they are receiving. The strategy now is to set a price that is comfortably above the cost floor and slightly below the value ceiling, then use real-world churn and upgrade data to calibrate the final number.

Interestingly, the market is signaling that AI users are more loyal when they pay more. Data from Venture Curator, which analyzed 3,500 companies, reveals a striking correlation between price points and customer retention. AI products priced at 250 dollars or more per month maintain a retention rate of approximately 70%. In contrast, products priced between 50 and 249 dollars see retention drop to 45%, and those priced under 50 dollars plummet to a mere 23%. This suggests that high-ticket AI tools are viewed as critical infrastructure or professional investments, whereas low-cost tools are treated as disposable experiments.

This realization is driving a shift toward outcome-based pricing, where the company charges for the result rather than the tool. The industry is currently experimenting with four primary agent-based models. The first is per-agent pricing, which positions the AI as a digital employee replacing a full-time human role. The second is per-action pricing, which is transparent but risks commoditization as AI capabilities become cheaper. The third is per-workflow pricing, charging for the completion of a complex sequence of tasks. The final, and perhaps most potent, is per-outcome pricing.

Intercom has already implemented this with its AI chatbot, Fin. Instead of charging a flat monthly fee for the bot, Fin charges 0.99 dollars per successfully resolved conversation. This aligns the company's incentives perfectly with the customer's: the customer only pays when the AI actually solves a problem, and Intercom captures value based on the efficiency of the resolution rather than the number of messages sent.

Ultimately, the era of the 90% margin is over for those building on LLMs. The new winners will not be those who can scale the fastest, but those who can most accurately map their pricing to the narrow corridor between infrastructure costs and delivered value.