Skip to main content
Zwei ASUS Ascent GX10 in dunklem Studio mit cyan und magenta Licht

Local LLMs at work: AI without the cloud, without the token bill

Published on September 14, 2026By Philipp Baldauf, CEO & Co-Founder

Plenty of teams want ChatGPT-level quality. They still cannot send contracts, customer data or internal knowledge to an API vendor. That is why we run local LLMs.

A local LLM is a language model that runs on hardware inside your own network. Prompts, files and answers never leave the building. There is no token bill from OpenAI or Anthropic. And the GDPR conversation gets calmer, because the data does not leave in the first place.

We did not just try this. We run it for ourselves and for clients.

Two ASUS Ascent GX10 units. DeepSeek V4 Flash. An interface that feels like a product, not a research project.

Local AI: a question to company knowledge, an answer with sources.

Two boxes on a desk. An AI system that does not need the cloud.

2

ASUS Ascent GX10 units in use

128 GB

Unified memory per unit

€0

API cost per token

What we built

The hardware is two ASUS Ascent GX10 units. Each box has the GB10 Grace Blackwell Superchip, 128 GB of coherent unified memory and up to 1 petaFLOP at FP4. It is 15 by 15 by 5 centimetres and plugs into a normal outlet. No server room. No rack. A desk.

On that hardware we run DeepSeek V4 Flash. Fast enough for daily work, strong enough for agents and long context.

In front of it sits the interface people actually use: chat, permissions, knowledge bases, agents, model choice. For the client that is not a setup project. It is an internal AI chat that can see company knowledge and works on day one.

We call it Ahoi AI.

What teams do with it

  • Ask policies, contracts, proposals and internal wikis, and get sources instead of guesses
  • Build agents for sales, support, HR or finance that only do what you allow
  • Switch between a local model and a cloud model when the task needs it
  • Keep knowledge in the house instead of inside a US prompt

How retrieval and knowledge bases work under the hood is in From long prompts to RAG. This piece is the full stack: hardware, model, interface, operations.

Ahoi AI Chat im MacBook
Ahoi AI on a MacBook: local model, agent, and sources from the knowledge base.
Ahoi AI Agenten im MacBook
Agents in Ahoi AI: their own knowledge and tone, local or another model.

How the stack runs

  1. Hardware in the building. The two ASUS Ascent GX10 units sit with us or on the client network. Data stays there.
  2. Model locally. DeepSeek V4 Flash covers daily work. We add other models, local or via API, when the task asks for it.
  3. Interface with knowledge and agents. Chat, files, agents, permissions. Ready for the people who will use it, not only for us.

Local versus a cloud API

Short version: the cloud is fast and strong. Local keeps the data in the building and drops the token bill. The comparison below is the rest.

What local does not magically fix

A local LLM is not always the best model in the world. Some jobs stay better in the cloud. That is why Ahoi AI can attach other models when you want that.

The boxes are not free either. They replace a monthly token bill with an investment that pays off once a team actually works with it.

And "local" alone does not make a clean organisation. Permissions, logs, sources, what the model may see: we set that up with you. Otherwise you just have a fast chat with no brakes.

Comparison of local LLMs versus cloud APIs
Local versus cloud: data, cost, control.

Common questions about local LLMs

If you want AI in the company without pouring every document into an API, this is the path we take ourselves and build for clients.

Hardware, model, interface, knowledge, agents, operations. A system that stays in the building.

Want AI that stays in the building?