Ten dollars a month for unlimited access to some of the largest language models on the planet. No metering, no surprise bills, just a flat fee and an API key that never stops working.
It's the kind of pitch that makes engineers either lean forward or reach for their calculators.
The platform is called sllm, and it appeared on Hacker News on April 5, 2026, with a model that challenges the conventional wisdom of AI infrastructure pricing. Instead of charging per token—the industry standard—sllm asks developers to join what it calls "cohorts," splitting the cost of a dedicated GPU with a handful of strangers. Think of it as a carpool for inference.
The idea isn't entirely new. Shared computing resources have existed since the mainframe era, and GPU timeshares have been around for years. But applying that model to language model APIs, at consumer-grade prices, with unlimited usage? That combination is rare enough to raise eyebrows.
And questions. Lots of them.
The Carpool Lane for LLMs
Here's how it works, at least in theory. You browse available cohorts on sllm's site, filtering by monthly price (ranging from $10 to $40), commitment period (one or three months), and advertised throughput—typically somewhere between 15 and 35 tokens per second per user. The model catalog includes some of the heavier hitters in open-weight AI: Llama 4 Scout, Qwen, GLM-5, DeepSeek.
You pick a cohort, your payment card gets authorized through Stripe, and you wait. The charge doesn't process until enough people sign up to fill the cohort and justify spinning up a dedicated node. If seven days pass without reaching capacity—a feature the team says is coming—the hold expires and you're free to try again elsewhere.
Once a cohort is live, you get an API key and access to the model for the duration of your subscription. When the term ends, you either join a new cohort or walk away. No penalties, no long-term contracts.
At least, that's the plan. When sllm launched, the site displayed a dashboard with filters and model names but no active cohorts to join. It was a storefront without inventory, waiting for enough interest to materialize.
"We're still early," wrote jrandolf, the username behind the Hacker News post, in response to one commenter. Early might be an understatement.
The Asterisk on 'Unlimited'
The word "unlimited" does a lot of work in sllm's pitch, perhaps more than it should. Yes, there are no token caps in the traditional sense—no moment when your API key stops working because you've exceeded a quota. But there are limits, just softer ones.
Each cohort lists a token rate, which functions as a performance target rather than a hard ceiling. The company frames this as optimization: continuous batching, efficient scheduling, the kind of engineering that lets multiple users share GPU cycles without stepping on each other's requests too badly.
Under the hood, sllm relies on vLLM, an open-source serving framework that's become popular for exactly this use case. It keeps model weights loaded in VRAM and processes multiple requests in the same forward pass, a technique that works well when usage is staggered. Average latency, according to jrandolf, runs under two seconds to first token. Worst case? Ten to thirty seconds.
The problem, as several Hacker News commenters pointed out with varying degrees of politeness, is concurrency. What happens when everyone in your cohort needs tokens at the same time?
One user ran the math. If a node delivers 3,000 tokens per second total and a cohort advertises 20 tokens per second across 465 slots, only about a third of the cohort can hit that rate simultaneously. The rest will wait. The calculations were rough—based on assumptions about cohort size and model throughput—but they illustrated a tension that sllm hasn't fully addressed.
Unlimited access, it turns out, depends on everyone using it sparingly.
The Fairness Problem

This is where the model gets interesting, or depending on your perspective, concerning. How does sllm handle contention? Is there rate limiting? Priority queues based on payment tier or time in queue? The FAQ doesn't say, and the company hasn't published service-level agreements or independent benchmarks.
Traffic routing, according to the site, happens through an "isolated proxy layer," and prompts are never logged. The infrastructure runs on what sllm describes as "dedicated GPU providers," though specific vendors aren't named. It's a vague enough description to invite skepticism, especially from developers who've been burned by opaque infrastructure before.
Several commenters raised the fairness question explicitly. If ten people in a cohort are building chatbots that spike during U.S. business hours and five others are running batch jobs overnight from Europe, the experience will be wildly different. One group gets their advertised token rate. The other gets whatever's left.
sllm's bet—unstated but implicit—is that usage will balance out. That cohorts will naturally mix time zones, use cases, and traffic patterns. That most developers will treat the service as a supplement rather than a primary workload.
Maybe. But it's a bet that needs real users to validate, and as of the launch, there weren't any.
A Crowded Field with a New Angle
sllm isn't entering a vacuum. OpenRouter aggregates models from dozens of providers, offering per-token pricing and the ability to route requests based on availability and cost. Vast.ai and similar platforms rent raw GPU capacity, letting teams run their own serving stack without abstraction layers. Replicate, Baseten, Modal—the landscape is full of companies trying to make inference easier, cheaper, or both.
What sllm offers, in theory, is predictability. A fixed monthly cost instead of a variable one. No surprise bills at the end of the month because your chatbot got popular or a contractor went rogue with the API key. For teams with consistent, high-volume usage, that predictability has value. But it only matters if the service actually delivers.
And right now, it's unclear whether it can. The model catalog is impressive on paper: billion-parameter models that would be expensive to run on per-token pricing. But without active cohorts or published performance data, there's no way to know if the economics work under real load—though the planned 7-day auto-cancel window for unfilled cohorts suggests performance data may emerge once testing begins. Early users will be beta testers, discovering whether a shared GPU can handle the traffic of multiple developers competing for cycles.
The Wait-and-See Phase

As of April 6, 2026, sllm is a concept with a dashboard. The infrastructure exists—jrandolf referenced time-to-first-token metrics and vLLM optimizations—but the business model is untested. The site is live, the pricing is posted, and the promise is bold. What's missing is the cohorts themselves, the actual communities of developers willing to bet $10 to $40 a month on shared access to a GPU they'll never see.
There's something appealing about the proposition. The idea that AI infrastructure could work more like a gym membership than a utility bill. Flat fee, unlimited use, as long as you're willing to share the equipment.
Whether that idea survives contact with real workloads—with developers who need tokens now, not when the queue clears—remains an open question. The kind of question that doesn't get answered with pitch decks and landing pages.
It gets answered when the first cohort fills, the first API key goes live, and someone tries to run a production workload on a GPU they're splitting with strangers.
That hasn't happened yet. But if it does, we'll know soon enough whether "unlimited" means what sllm says it does—or just sounds like it should.
