Everything you'd expect from the big AI labs, running the best open models on private compute. One key, one API, one line to switch. Nothing you send is stored, nothing is trained on, and not even we can read it.
Three models, each named for the job it does. Fetch for the calls nobody reads, Do for the work you watch, Decide for the work you hand over. Every model ships with its full context*. Pick the pace only when something is waiting on it.
| Priority |
|---|
fetch@fast
click any price to try it, then paste the string wherever your tool asks for a model
* Fair-use context cap, per hour. Most tools cache by default, so it rarely applies.
It speaks the language your tools already do. Point your editor, your agent or your pipeline here and carry on exactly as you were. Nothing to rebuild, nothing new to learn.
Change your mind mid‑task, as often as you like. Whatever you are not using costs you nothing, because nothing of yours is running.
# one key, every tier base_url: "https://api.lararia.ai/v1" api_key: "opt_live_••••••••••••" model: do@fast # ↑ change this, and nothing else
Give everyone a key. One key runs as many agents as you like. The model string carries the model, the priority, and nothing else to configure. You pay for what is running right now. That's the whole idea.
Most AI services ask you to trust a privacy policy. Private compute removes the need to. Your prompt is encrypted on your machine and only ever decrypted inside the sealed chip that answers it. It is used for your request, then wiped. Not stored. Not readable by us. Not trained on. On every request, by default.
Your machine encrypts the prompt before it leaves, and it alone can open the reply.
Plain textEverything in transit is TLS-encrypted. Anyone on the wire sees scrambled bytes.
CiphertextWe see the address on the envelope: which key, which model. Never the letter inside.
CiphertextGPU memory is encrypted by the hardware itself. The key is born inside the chip and never leaves it, so even someone at the machine, with admin or physical access, reads noise.
NothingThe one other place the prompt is plain text: inside the sealed chip, while it answers. Wiped the moment the last token streams out.
Plain text, sealedThis is built for the work a business can't paste into a chatbot: client files, contracts, patient records, source code, customer data. The question "can we send this?" goes away, because nobody in the middle can read it: not us, not the datacenter, not anyone on the wire.
And it is not an enterprise add-on you negotiate for. Every key, every model, every request runs this way. Your data passes through sealed memory and is gone: nothing written down, nothing kept, nothing learned from.
And if geography matters to you: it all runs in Europe, in Paris and Frankfurt, under European law.
Card, key, first request. The best open models, run the way they deserve to be: simple for the developer, private for the business.