
Local LLMs at work: AI without the cloud, without the token bill
Plenty of teams want ChatGPT-level quality. They still cannot send contracts, customer data or internal knowledge to an API vendor. That is why we run local LLMs.
A local LLM is a language model that runs on hardware inside your own network. Prompts, files and answers never leave the building. There is no token bill from OpenAI or Anthropic. And the GDPR conversation gets calmer, because the data does not leave in the first place.
We did not just try this. We run it for ourselves and for clients.
Two ASUS Ascent GX10 units. DeepSeek V4 Flash. An interface that feels like a product, not a research project.
Two boxes on a desk. An AI system that does not need the cloud.
ASUS Ascent GX10 units in use
Unified memory per unit
API cost per token
What we built
The hardware is two ASUS Ascent GX10 units. Each box has the GB10 Grace Blackwell Superchip, 128 GB of coherent unified memory and up to 1 petaFLOP at FP4. It is 15 by 15 by 5 centimetres and plugs into a normal outlet. No server room. No rack. A desk.
On that hardware we run DeepSeek V4 Flash. Fast enough for daily work, strong enough for agents and long context.
In front of it sits the interface people actually use: chat, permissions, knowledge bases, agents, model choice. For the client that is not a setup project. It is an internal AI chat that can see company knowledge and works on day one.
We call it Ahoi AI.
What teams do with it
- Ask policies, contracts, proposals and internal wikis, and get sources instead of guesses
- Build agents for sales, support, HR or finance that only do what you allow
- Switch between a local model and a cloud model when the task needs it
- Keep knowledge in the house instead of inside a US prompt
How retrieval and knowledge bases work under the hood is in From long prompts to RAG. This piece is the full stack: hardware, model, interface, operations.


How the stack runs
- Hardware in the building. The two ASUS Ascent GX10 units sit with us or on the client network. Data stays there.
- Model locally. DeepSeek V4 Flash covers daily work. We add other models, local or via API, when the task asks for it.
- Interface with knowledge and agents. Chat, files, agents, permissions. Ready for the people who will use it, not only for us.
Local versus a cloud API
Short version: the cloud is fast and strong. Local keeps the data in the building and drops the token bill. The comparison below is the rest.
What local does not magically fix
A local LLM is not always the best model in the world. Some jobs stay better in the cloud. That is why Ahoi AI can attach other models when you want that.
The boxes are not free either. They replace a monthly token bill with an investment that pays off once a team actually works with it.
And "local" alone does not make a clean organisation. Permissions, logs, sources, what the model may see: we set that up with you. Otherwise you just have a fast chat with no brakes.

Common questions about local LLMs
A local LLM is a language model that runs on hardware inside your own network. Prompts, files and answers never go to a cloud vendor. That is how we run DeepSeek V4 Flash on two ASUS Ascent GX10 units.
They fix the data problem at the root: content does not leave the building. That does not replace the rest of your setup (permissions, logs, processing internally). We put that in place with you.
You pay hardware and setup once, then no token bill. Two ASUS Ascent GX10 units replace ongoing API spend once a team actually works with the system. Operations and knowledge care are part of the job, and part of what we deliver.
Yes for chat, agents and models in the DeepSeek V4 Flash class. One box has 128 GB of unified memory and up to 1 petaFLOP at FP4. We run two so load and spare capacity line up. Two GX10 units can be linked when models get larger.
Yes. Ahoi AI can attach other models too. Local is the default for internal knowledge. Cloud stays an option when a job needs a stronger model and the data allows it.
If you want AI in the company without pouring every document into an API, this is the path we take ourselves and build for clients.
Hardware, model, interface, knowledge, agents, operations. A system that stays in the building.