When we talk to people who run firms about a private AI appliance, the second question is nearly always the same. The first is what it does. The second is how many of us can use it.
It's a fair question and it's hard to get a straight answer to. Cloud AI providers never have to answer it, because their capacity is somebody else's problem. Most people selling AI that runs on your own premises quote a number without saying what it was measured on, which makes the number close to useless.
So we measured ours carefully, and wrote down the conditions. The figures below are smaller than a sales page would choose. We think that's the right way round.
The short answer
One Kubrius Core appliance handled ten people waiting on answers about long documents at the same time, with the answer starting to appear within ten seconds for 95 in every 100 requests.
A long document here means roughly 4,800 words going in, which is about ten pages. At twelve people the wait crept past ten seconds, so ten is where we draw the line.
Shorter work gets a stricter standard. Someone asking a quick question shouldn't wait ten seconds for the reply to start, because at that point most people assume the thing has hung. So we hold quick questions to two seconds and everyday document questions to five.
| Kind of work | Answer starts within | People at once |
|---|---|---|
| Quick question about 180 words in | 2 seconds | At least 10, fewer than 20 |
| Everyday document about 1,200 words in | 5 seconds | At least 10, fewer than 16 |
| Long document about 4,800 words in | 10 seconds | 10 |
The first two rows are ranges because we haven't yet tested every step between them. We'd rather give you a range we can stand behind than a single number we'd have to guess.
People waiting, not people employed
Ten sounds small, so it's worth being clear about what it counts.
It counts people who have asked something and are waiting for the answer at that exact moment. It doesn't count staff, licences or logins. Someone using AI well spends most of their time reading what came back, checking it against the file and deciding what to ask next. While they do that, the appliance is free for somebody else.
It also leaves out everything that never touches the appliance. Kubrius Core has two routes. Everyday work that isn't confidential goes out to an approved public model, much as it would today. Only confidential work stays on the box. The ten is shared among the confidential requests, not among everything your firm does with AI.
What we measured
The appliance is an NVIDIA GB10 system with 128 GB of memory, small enough to sit on a shelf in a comms cupboard. It runs Qwen3.6-35B-A3B, an open-weight model, entirely on the box.
We sent it requests in the three shapes in the table. Each simulated person asked three questions one after another, and we kept adding people until answers started arriving too slowly. The rest of the Kubrius software stayed running the whole time, so the model was competing for memory the way it would in real use.
Two things we learned along the way are worth passing on, because each one moved the number a long way.
Our first test flattered us. It reused the same questions, and the machine recognised work it had already done and skipped part of it. Real questions are never exact repeats. Once every request had its own fresh text, the figures went down, and the lower figures are the ones on this page.
Then we found a problem in the other direction. The software that runs the model has a setting for how many requests it will work on at once, and ours was set to five. Everyone after the fifth was simply queuing. Raising that one setting took the long-document figure from six people to ten, on the same hardware.
We also tried a newer model that scores better on published benchmarks. On our appliance it managed one person on long documents within the same ten seconds, against ten for the model we kept. A model can be better on paper and still be the wrong choice for a small machine, and you only find out by running it.
Working it through for your firm
This is how we'd size a firm, with the assumptions out in the open so you can swap in your own.
Take a firm of 120 people. Suppose that at the busiest moment of a normal day, one person in ten is waiting on an AI answer. That's twelve people. Suppose a third of that work is confidential and stays on the appliance. That's four. The tested limit for the heaviest kind of work is ten, so there's plenty of room.
Now take 250 people with the same assumptions. Twenty-five are waiting, and about eight of them are on the appliance. That's still inside ten, but closer than we'd like. We keep about a quarter spare for bad days, which puts the comfortable limit at seven. At that size we'd want to look at how your people actually use AI before promising anything.
Any of those assumptions could be wrong for your firm. One in ten might be one in twenty. A third confidential might be half. Busy moments also tend to be busy for everyone at once: month end, a big completion, a deadline the whole team is working to. That's exactly when averages stop holding. Treat the arithmetic as the start of a conversation. It isn't a promise.
What this doesn't prove yet
We'd rather you read this from us than work it out for yourself later.
- We sent requests straight to the model. A real request also searches your documents, ranks what it finds, checks the policy and writes an audit record. Each of those adds time, so the wait a person actually sees will be longer than the figures above.
- The answers were short, around 70 words. Real answers are often several times longer. That doesn't change how quickly an answer starts, but each person holds their place for longer.
- The text was written for the test. It was the right length, but it wasn't anybody's contracts, spreadsheets or scanned letters.
- It was one appliance, tested over hours rather than days.
- Kubrius Core is in beta. It runs on our own appliance today and has never run inside anybody else's building. The first firms to use it will be the ones who turn these estimates into real numbers.
Five questions to ask any vendor
Whoever you end up talking to, including us, these will tell you how much a capacity figure is worth.
- How many people waiting at once, and how long before the answer starts? A number of users with no time attached doesn't tell you anything.
- How much text was going in? Ten people asking one-line questions is a very different claim from ten people sending in a contract.
- Was the rest of the system running during the test? A model measured on its own, on an otherwise idle machine, will always look better than it performs.
- Were the test questions all different? If they repeated, the result is flattering.
- What happens past the limit? Ours slows down rather than failing. At twelve people on long documents the answers took longer to start, but every request still completed.
If you'd like our full test report, it goes into the method and the limits in more detail than there's room for here. Use the form below, mention the test report, and we'll send it over.
Frequently Asked Questions
Why is the long-document limit ten?
Most of the time goes on reading what you sent before any of the answer is written. A 4,800-word document is a lot of reading for one small machine, and ten of them at once is where the wait passes ten seconds.
Does the non-confidential work count towards the ten?
No. That goes out to an approved public model and doesn't use the appliance at all.
Would a second appliance double it?
Probably close to it, but we haven't tested two together yet, so we won't claim it.
Why not use a bigger or newer model?
We tested one. It scored better on published benchmarks and served one person where ours served ten. On a machine this size, keeping up with several people at once matters more than a few points on a leaderboard.
Is ten a guarantee?
No. It's what we measured under the conditions on this page. Your documents and your people's habits will move it, which is why we'd measure at your firm before anyone relies on it.