Online qwen3.8-27b : 31m·qwen3.8-flash : 31m
OpenAI-compatible · zero retention · flat per-seat pricing

Inference for open-weight models,
at interactive speed. beta

sevenk.ai serves open-weight models through one OpenAI-compatible endpoint. Point any agent, IDE, or SDK at it and stream.

https://sevenk.ai/v1OpenAI-compatible base URL

Currently running

These cards show exactly what is routable right now. The same catalogue is machine-readable at GET /models.

Qwen3.8 27b

qwen3.8-27b
27B262k contextNVFP4ToolsReasoningVisionChatCompletions
There is no per-token cost. Decoding is included with your seat.

Qwen3.8 Flash

qwen3.8-flash
125B A6B179k contextNVFP4ToolsReasoningVisionChatCompletions
There is no per-token cost. Decoding is included with your seat.

Vision input is experimental; assume text for anything you depend on. Served models may change as capacity is upgraded.

Connect your agent

Add sevenk.ai as a provider anywhere an OpenAI-compatible endpoint will do. The base URL is https://sevenk.ai/v1, and your key goes in the Authorization: Bearer header. Pick your harness below:

Register sevenk.ai as a custom provider in ~/.config/opencode/opencode.json, then authenticate once via /connect (choose sevenk.ai → key) or by setting SEVENK_API_KEY on the OpenCode server process. Models are discovered from the endpoint automatically.

{
  "$schema": "https://opencode.ai/config.json",
  "model": "sevenk.ai/qwen3.8-27b",
  "providers": {
    "sevenk.ai": {
      "name": "sevenk.ai",
      "package": "@opencode/ai/providers/openai-compatible",
      "settings": {
        "baseURL": "https://sevenk.ai/v1",
        "apiKey": "{env:SEVENK_API_KEY}"
      }
    }
  }
}

Performance

These are estimates based on measurements of active performance. The curves are anchored to live single-seat runs measured on this service; the first of those were captured in an earlier measurement session. Above the single-seat rate, each concurrency level is modelled from burst observations rather than sustained load. Capacity is shared: your rate is your fair share of whatever the service is doing at the moment, ordered by the fair-use governor.

Qwen3.8 27b
Estimated top output speed by concurrency (qwen3.8-27b)
0 250 500 750 1000 205372548702812 concurrent requests tokens / second
concurrent reqs124816
total tokens/s205372548702812
per-seat tokens/s2051861378851
Qwen3.8 Flash
Estimated top output speed by concurrency (qwen3.8-flash)
0 125 250 375 500 88152224281322 concurrent requests tokens / second
concurrent reqs124816
total tokens/s88152224281322
per-seat tokens/s8876563520

On the roadmap

We are actively accelerating inference. The steps below are in motion, not intentions, and each one adds hardware that makes more of the service available: storage is arriving now, more memory and a second node follow. New models appear here, and at GET /models, the moment they become routable.

  1. Demand-dependent

    Live capacity view

    We will publish tokens/minute and queue depth, kept in memory only, alongside the model list. It can arrive before any hardware step below, whenever usage makes the numbers worth watching.

  2. In progress

    More storage

    More storage is on its way to the serving hardware. It comes before the memory upgrade and the resilience work below.

  3. Next

    More server memory

    Additional memory arrives once the storage is in place. It widens how much the service can hold and run at once.

  4. After memory

    Higher availability

    Battery backup and a second node follow, and when they land, downtime and maintenance fall with them.

Next in the model pipeline

DeepSeek V4.1

Long-context reasoning
Planned
It is next in the queue for a serving slot.

GLM 5.3 Flash

Fast instruct for agent work
Planned
It is queued behind DeepSeek V4.1.

Get started

beta The whole service is in beta, pricing and policies included. One seat is one person and one key.

Request an API key

We review each request personally, then reply with your key and connection details. You do not need an account or a card.

We keep only your name and email, and only until the key is issued.
When your key arrives, export SEVENK_API_KEY=sk-gw-… and pick a recipe above.