Inference for open-weight models,
at interactive speed. beta
sevenk.ai serves open-weight models through one OpenAI-compatible endpoint. Point any agent, IDE, or SDK at it and stream.
https://sevenk.ai/v1OpenAI-compatible base URL- Zero retention. We never persist prompts, completions, files, or transcripts. Retention is impossible by construction, not just disallowed by policy.
- Interactive speed. A single seat decodes above 200 tokens/second when the service is quiet. How the rate moves as it fills is charted in Performance.
- Flat per-seat pricing. A seat costs $25/month and includes unlimited daily tokens. There is no per-token metering and no overage bill, and you can cancel anytime.
Currently running
These cards show exactly what is routable right now. The same catalogue is machine-readable at GET /models.
Qwen3.8 Flash
Vision input is experimental; assume text for anything you depend on. Served models may change as capacity is upgraded.
Connect your agent
Add sevenk.ai as a provider anywhere an OpenAI-compatible endpoint will do. The base URL is https://sevenk.ai/v1, and your key goes in the Authorization: Bearer header. Pick your harness below:
Register sevenk.ai as a custom provider in ~/.config/opencode/opencode.json, then authenticate once via /connect (choose sevenk.ai → key) or by setting SEVENK_API_KEY on the OpenCode server process. Models are discovered from the endpoint automatically.
{
"$schema": "https://opencode.ai/config.json",
"model": "sevenk.ai/qwen3.8-27b",
"providers": {
"sevenk.ai": {
"name": "sevenk.ai",
"package": "@opencode/ai/providers/openai-compatible",
"settings": {
"baseURL": "https://sevenk.ai/v1",
"apiKey": "{env:SEVENK_API_KEY}"
}
}
}
}
pi does no server-side discovery, so list the models in ~/.pi/agent/models.json. contextWindow comes from GET /v1/models; maxTokens is a client-side cap you choose.
// ~/.pi/agent/models.json
{
"providers": {
"sevenk": {
"baseUrl": "https://sevenk.ai/v1",
"api": "openai-completions",
"apiKey": "sk-gw-…",
"compat": { "supportsDeveloperRole": false },
"models": [
{ "id": "qwen3.8-27b", "reasoning": true, "input": ["text","image"], "contextWindow": 262144, "maxTokens": 32768 },
{ "id": "qwen3.8-flash", "reasoning": true, "input": ["text","image"], "contextWindow": 179200, "maxTokens": 32768 }
]
}
}
}
In Cline, Roo Code or Kilo Code pick the OpenAI Compatible provider (not the built-in Ollama or LM Studio ones) so inference rides /v1/chat/completions and the model list comes from /v1/models.
Provider OpenAI Compatible
Base URL https://sevenk.ai/v1
API key sk-gw-…
Model qwen3.8-27b
Enter the model id by hand if it does not appear in the picker.
- Settings → Models → Override OpenAI Base URL =
https://sevenk.ai/v1 - Add your API key, then Add Custom Model for each model id and press Verify
- Select the model for Cursor Agent (or chat) in the model picker
Do not prefix model ids with claude-*: Cursor rewrites those to Anthropic's wire format.
Codex CLI does no discovery, so pin the model and point the provider at the gateway in ~/.codex/config.toml, then export SEVENK_API_KEY.
# ~/.codex/config.toml
model = "qwen3.8-27b"
model_provider = "sevenk"
[model_providers.sevenk]
name = "sevenk.ai"
base_url = "https://sevenk.ai/v1"
wire_api = "responses"
env_key = "SEVENK_API_KEY"
Anything that speaks OpenAI works: SDKs, Aider, LobeChat, n8n, scripts. Stream for long generations; errors keep the OpenAI shape.
curl https://sevenk.ai/v1/chat/completions \
-H "Authorization: Bearer sk-gw-…" \
-H "Content-Type: application/json" \
-d '{
"model": "qwen3.8-27b",
"stream": true,
"messages": [{ "role": "user", "content": "hello" }]
}'
import OpenAI from "openai";
const client = new OpenAI({
baseURL: "https://sevenk.ai/v1",
apiKey: process.env.SEVENK_API_KEY,
});
const stream = await client.chat.completions.create({
model: "qwen3.8-27b",
stream: true,
messages: [{ role: "user", content: "hello" }],
});
Performance
These are estimates based on measurements of active performance. The curves are anchored to live single-seat runs measured on this service; the first of those were captured in an earlier measurement session. Above the single-seat rate, each concurrency level is modelled from burst observations rather than sustained load. Capacity is shared: your rate is your fair share of whatever the service is doing at the moment, ordered by the fair-use governor.
| concurrent reqs | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|
| total tokens/s | 205 | 372 | 548 | 702 | 812 |
| per-seat tokens/s | 205 | 186 | 137 | 88 | 51 |
| concurrent reqs | 1 | 2 | 4 | 8 | 16 |
|---|---|---|---|---|---|
| total tokens/s | 88 | 152 | 224 | 281 | 322 |
| per-seat tokens/s | 88 | 76 | 56 | 35 | 20 |
On the roadmap
We are actively accelerating inference. The steps below are in motion, not intentions, and each one adds hardware that makes more of the service available: storage is arriving now, more memory and a second node follow. New models appear here, and at GET /models, the moment they become routable.
- Demand-dependent
Live capacity view
We will publish tokens/minute and queue depth, kept in memory only, alongside the model list. It can arrive before any hardware step below, whenever usage makes the numbers worth watching.
- In progress
More storage
More storage is on its way to the serving hardware. It comes before the memory upgrade and the resilience work below.
- Next
More server memory
Additional memory arrives once the storage is in place. It widens how much the service can hold and run at once.
- After memory
Higher availability
Battery backup and a second node follow, and when they land, downtime and maintenance fall with them.
Next in the model pipeline
DeepSeek V4.1
GLM 5.3 Flash
Get started
beta The whole service is in beta, pricing and policies included. One seat is one person and one key.
Request an API key
We review each request personally, then reply with your key and connection details. You do not need an account or a card.