Private compute, on by default · on every request

The best models, open.
Private by default.

Everything you'd expect from the big AI labs, running the best open models on private compute. One key, one API, one line to switch. Nothing you send is stored, nothing is trained on, and not even we can read it.

Not even we can read your prompts Nothing stored, nothing trained on Open models, yours to take elsewhere Nothing billed when idle Runs in Europe

Always the right model.

Three models, each named for the job it does. Fetch for the calls nobody reads, Do for the work you watch, Decide for the work you hand over. Every model ships with its full context*. Pick the pace only when something is waiting on it.

Best fit for
Priority
Model string fetch@fast click any price to try it, then paste the string wherever your tool asks for a model

* Fair-use context cap, per hour. Most tools cache by default, so it rarely applies.

Switching models is changing a word.

It speaks the language your tools already do. Point your editor, your agent or your pipeline here and carry on exactly as you were. Nothing to rebuild, nothing new to learn.

Change your mind mid‑task, as often as you like. Whatever you are not using costs you nothing, because nothing of yours is running.

~/.config/agent.yaml
# one key, every tier
base_url: "https://api.lararia.ai/v1"
api_key:  "opt_live_••••••••••••"

model: do@fast
#      ↑ change this, and nothing else

You are never charged for existing.

Give everyone a key. One key runs as many agents as you like. The model string carries the model, the priority, and nothing else to configure. You pay for what is running right now. That's the whole idea.

Nothing running Live: keys start and stop on their own
Flat line. The keys all still exist. None of them is billing.
Org spend, this month
€63.1500
Combined rate
€0.00 /h
Agents running
0 now
KeyRunning nowThis monthDelete
6 keys · pay only for the ones running
Both halves of the string are names, not numbers. The model says what the work is: fetch, do, decide. The pace says who is waiting: normal for a person, fast for a machine. Every model comes with its full context included. What you're billed for is simple: the model actually thinking, nothing more. Tool calls and waiting on you don't count.

Not even we can read what you send.

Most AI services ask you to trust a privacy policy. Private compute removes the need to. Your prompt is encrypted on your machine and only ever decrypted inside the sealed chip that answers it. It is used for your request, then wiped. Not stored. Not readable by us. Not trained on. On every request, by default.

This is built for the work a business can't paste into a chatbot: client files, contracts, patient records, source code, customer data. The question "can we send this?" goes away, because nobody in the middle can read it: not us, not the datacenter, not anyone on the wire.

And it is not an enterprise add-on you negotiate for. Every key, every model, every request runs this way. Your data passes through sealed memory and is gone: nothing written down, nothing kept, nothing learned from.

Private compute on every request, not a paid tier
Never used for training, ever
Zero retention, on by default
DPA and Article 28 terms, signable online
Open weights only, portable the day you leave

And if geography matters to you: it all runs in Europe, in Paris and Frankfurt, under European law.

This is what running it should feel like.

Card, key, first request. The best open models, run the way they deserve to be: simple for the developer, private for the business.