beta sevenk.ai runs in beta, and that covers this document too: prices and terms change as capacity grows. The version block below states what is in force today.
Service Level Agreement
Version 1.1 · Effective 2 October 2026 · Applies to all paid seats
This SLA is written for the service sevenk.ai actually is today: one deployment, one operator, no battery backup, no failover target yet. Rather than pretend at five nines, it defines an honest maintenance window, commits to an availability number the current hardware can actually hit, and pays credits when we miss it. Each resilience improvement on the roadmap (§9) raises the number it was written against.
1. Definitions
- Service: the inference gateway accepting requests from a valid seat key and returning completions or explicit errors.
- Available: the gateway accepts authenticated requests and returns either a completion or a defined HTTP error within the request timeout. Model quality is never part of availability.
- Downtime: minutes when the Service is not Available to seats.
- Maintenance Window: the scheduled period defined in §2, excluded from availability.
- Service Month: a calendar month, US Eastern.
- Availability: (total minutes − maintenance − downtime) ÷ (total minutes − maintenance) × 100.
2. Maintenance Windows
A standing window runs nightly, 9:00 PM to 1:00 AM US Eastern, for up to four hours; most nights expect three to four hours of interruption. Work inside that window needs no notice. Work outside it is announced at least 24 hours ahead where practicable and counts against the window quota rather than as downtime; unannounced substantial out-of-window maintenance counts as downtime.
Because there is no battery backup, a power interruption stops the service immediately. The restart target is within 30 minutes of power returning, and such interruptions count as downtime except as §6 provides.
3. Availability Commitment
Measured per Service Month, excluding maintenance:
| Monthly availability (ex-maintenance) | Status |
|---|---|
| ≥ 98.0% | Commitment met (internally we target ≥ 99%) |
| < 98.0% | Service credits apply (§4) |
Why 98% and not 99.9%: the serving hardware has no battery backup or failover yet, so grid flicker, driver faults and human repair time are unmitigated single points of failure. 98% ex-maintenance means up to roughly 12.4 hours of unplanned downtime in a 30-day month. We commit to the number we can actually hit, and we raise this SLA with a published version bump whenever a real resilience improvement ships (§9), instead of promising it first.
4. Service Credits: the Sole Remedy
If §3 is missed in a Service Month, the seat holder may claim:
| Availability (ex-maintenance) | Credit (% of that seat's monthly fee) |
|---|---|
| < 98.0% | 10% |
| < 95.0% | 25% |
| < 90.0% | 50% |
| Unusable for ≥ 7 consecutive days | 100% of the affected days' prorated fee |
Credits are applied to the next invoice, or refunded if the seat is cancelled; cumulative credits are capped at one month's fee and are the sole and exclusive remedy for availability failures. Submit a claim within 15 days of month end. Because neither side has server logs (Policy §3), claims are judged from the public status notice and your own client-side evidence; timestamps of failed requests are enough. We do not dispute a claim the status notice cannot contradict.
5. Performance: Disclosed, Not Committed
Speed is a shared, capacity-limited resource, and only §3 is contractual. A single seat on an otherwise idle service has been observed decoding above 200 tokens per second, prompt-dependent; that burst is not sustained while other seats are generating, and per-seat rates fall as concurrency rises. Typical shared rates are not committed to.
The fair-use governor (Pricing §4) keeps one generation in flight per seat, queues the rest, and answers a full queue with 429 and a Retry-After rather than degrading everyone silently. Time-to-first-token is deliberately favored over aggregate throughput: queue fairness and quick starts beat a marginally higher tokens-per-second ceiling.
6. What Does Not Count as Downtime
- Maintenance-window activity (§2) and announced out-of-window work.
- Utility grid outages longer than one hour, treated as force majeure. Brief blips and brownouts that restart the serving hardware do count: there is no UPS yet, and adding one is on the roadmap (§9), so we own the outages it causes.
- Failures of your own network, device or client software.
- Rejections under the fair-use limits (§5), invalid keys, or suspension under Policy §5.
- Experimental capabilities: vision input is excluded from this SLA altogether and may error or disappear at any time.
- Force majeure: fire, storm, theft, ISP termination, grid collapse, and their ilk.
7. Incident Response
The operator is one person, so the response times below are honest rather than aspirational.
| Severity | Definition | First response |
|---|---|---|
| P1 (service down) | Gateway unreachable for all seats | ≤ 4h (≤ 1h inside the maintenance window) |
| P2 (degraded) | Severe slowdowns, high error rates, model faults | ≤ 12h |
| P3 (question or request) | Billing, questions, feature asks | ≤ 3 business days |
The public status notice is the single source of truth for incident history, and it records system events only, never request content. Requests to recover lost data are always refused: nothing exists to recover.
8. Changes to This SLA
Material reductions take effect only after 30 days' notice, with the right to cancel without penalty. Improvements, a higher commitment after battery backup or redundancy ships, take effect on their published version bump. This SLA does not expire while your seat is continuously active.
9. The Roadmap, Tied to This SLA
None of the following is contractual, but each one buys a stricter promise: battery backup removes flicker-related downtime and lifts the commitment above 99%; a second node allows a maintenance-light tier; independent power lets us narrow or drop the grid-outage exclusion in §6. When any of them ships, this SLA is revised upward, not reinterpreted.