A model that keeps the whole problem in view.
Nexith Core is our flagship model, trained from the ground up for reasoning, writing, and code. It holds a 1M-token context, so it works over whole documents and codebases at once instead of fragments — and it reads images alongside text.
Semantic cache hits save up to 90% on input tokens. No credit card required to start.
Prompts, documents, and code — up to 1M tokens in a single request.
Read images alongside text in the same message.
Streamed token-by-token over SSE, up to 128K per response.
Define tools in JSON schema; the model calls them with typed arguments.
Constrain responses to a schema so you get valid JSON every time.
Persist context across sessions, scoped per user and per app.
Nexith Core is our flagship model, trained from the ground up for reasoning, writing, and code. It holds a 1M-token context, so it works over whole documents and codebases at once instead of fragments — and it reads images alongside text.
It runs on our own sovereign GPU cluster behind a clean, OpenAI-compatible API — with streaming everywhere, a semantic cache that makes repeated work instant, and memory that carries across sessions.
What makes it good
Serious coding
Writes, reads, and refactors code across languages — and knows when a task is actually done.
Long-horizon tasks
Breaks a goal into steps, tracks its own budget, and drives to a clean finish instead of spinning.
Persistent memory
Remembers across sessions — scoped per user and per app, so it picks up where you left off.
Semantic cache
Recognizes repeated work and answers instantly at up to 90% off, with no repeated compute.
