Topobius · Decision Intelligence Box

Expert-grade judgment and decisions,
in one desktop box.

Domain post-training, frontier-class local models, and a multi-stream agent runtime in one box. Real-time professional reasoning, with data that never leaves your domain. Desktop-class hardware: power and ethernet, and it stands watch — no machine room required.

Topobius Gradientron — desktop decision intelligence box
TOPOBIUS · AI HUB
0.24 s
Business context ready
Thunder-KV · model & runtime co-rebuilt
<0.3 s
First token
The domain plan speaks instantly
16 streams
Concurrent agents
Isolated contexts · zero switching cost
50–100 chars/s
Real-time professional reasoning
≈10× human thought speed

Intelligence and data, under local control — nothing about the operation leaves the site.

Wherever there are people, business and private data, intelligence should not live in someone else's rack.

Frontier capabilitycapability
On-premise has long been equated with "capability compromise" — small models only do intent recognition and chit-chat.
The built-in model passes official benchmarks: several capabilities exceed online flagships, and the overall profile sits in the current frontier range.
Data sovereigntysovereignty
Public-cloud routes: identity, accounts and business data leave the premises — compliance review at every step.
Reasoning, orchestration and action instructions all close the loop locally. Data stays in-domain.
Deterministic experienceexperience
Public links fluctuate — like driving across Beijing at rush hour, efficiency is never guaranteed.
LAN RTT under 50ms, no degradation offline, a fixed monthly electricity bill.
Build costcost
Per-terminal deployment: N devices, N maintenance, and new hardware everywhere for AI features.
One hub serves 8–16 terminals; existing devices live on as thin clients.
Frontier-class · measured capability

A model in a local box, competing on the same track as online flagships.

Official model-card public benchmarks: several capabilities exceed current online flagships; the overall profile sits in the frontier range — local deployment no longer means compromise.

BenchmarkAI-hub built-in modelClaude Opus 4.8GPT-5.6 SolDeepSeek V4 Flash
SWE-bench Pro · real code repair61.769.264.656.0
OSWorld-Verified · desktop ops84.383.483.2
IFBench · instruction following79.562.272.7
GPQA Diamond · science reasoning89.293.694.190.8
Terminal-Bench 2.1 · terminal coding73.074.688.882.7

Desktop operation and instruction-following lead both overseas flagships (OSWorld 84.3 vs 83.4 / 83.2; IFBench 79.5 vs 62.2 / 72.7); the GPQA science gap with each flagship is under 5 points; LiveCodeBench v6 90.3, OmniDocBench 1.5 91.1 (most flagships publish no comparable figures) — the overall profile sits in the current frontier range.

84.3OSWorld desktop ops · ahead of overseas flagships
79.5IFBench instruction following · ahead
90.3LiveCodeBench v6 · new-problem coding

Sources: official model cards & system cards of each model (vendor-reported, August 2026 snapshot); DeepSeek V4 Flash from its official card and third-party evaluation compendia.

Three layers, delivered as one

A machine built for business intelligence from the transistor up — not an inference framework with bolted-on agents.

Operators are built for the model, the model is tuned for the business, and the runtime hands every terminal's request to the right model. The layers mesh, and it works on arrival at the site.

Access layeraccess
sensorscameras · lidar · sensors · gaugesactuncrewed equipment · industrial automation under unified controlwssstreaming interaction for the command wallapibusiness-system integration
Agent runtimeagent runtime
coremulti-stream agent concurrency · isolated contextsorgenrolled against the org chartplanmulti-step task planningtoollocal tool & equipment invocationroutecomplexity-aware routing
Purpose-built model poolmodel pool
flagshipfrontier-class local model · domain post-trainingfastintent recognition · FAQ · slot extractionchatmulti-turn dialogue · tool orchestrationdeepcomplex reasoning · long-document understanding
Custom inference engineinference engine
opshardware-aware operators · full-chain fusionmemexplicit memory management · weight pinningschedcontinuous batching · hybrid schedulingpindedicated inference-core scheduling
Thunder-KV · model & runtime rebuilt together

Other systems are still loading.
Yours starts the moment it hears.

Domain protocols, playbooks, tool interfaces and expert knowledge are baked into the device at model-customization time. Ready at power-on, shared across sessions, zero switching cost. Co-innovation across silicon operators, model restructuring and agent interaction — all you feel is the speed.

50–100s
Generic approach — the prefill pain of local compute
0.24s
Thunder-KV — 50k-token business prefix ready, first token in under 0.3s
THUNDER-KV ≈ 200–400× faster business-context readiness
Generic approach50–100 s
Thunder-KV0.24 s
Reasoning & agents

Thinks like an expert. Moves like a hand on the floor.

reason

Real-time professional reasoning

A frontier-class local model, post-trained on industry and specialist knowledge. 50–100 chars/second of sustained output — roughly ten times human thought speed. Complex reasoning, long-document understanding, anomaly assessment, with domain-expert-grade decision proposals.

agent ×16

Multi-stream agent concurrency

Sixteen independent agents run in parallel on one machine, enrolled against the org chart with fully isolated contexts. Sensing, assessment, planning and instruction each do their own job; a single-instance failure never spreads.

plan + tool

Multi-step planning & tool use

Intent recognition, task decomposition, local API and equipment invocation, result validation, plan assembly — one chain completed inside the inference pipeline.

rag

Local knowledge augmentation

Business documents and product FAQs hot-update; retrieved results merge into context automatically, no restarts.

route

Complexity-aware routing

A lightweight model acts as gateway, judging request difficulty — simple intents answered instantly, complex reasoning handed to the deep tier. Compute spent where it counts.

offline

No degradation offline

When the internet drops, core reasoning and agent services keep running. The outside network is an update channel, not a lifeline.

audit

Full traceability

Every instruction and every disposition is traceable, replayable and auditable — the action chain closes completely.

The domain decision loop

The machine proposes. The person decides. The equipment executes.

The same engineering path as autonomous driving: a four-step closed loop of learn, forge, connect, decide. Every use makes the system stronger.

learn · knowledge injection

Literature, protocols, wargaming material, case files and equipment manuals — all domain material becomes training data. Decades of accumulated expertise becomes machine-usable capability for the first time.

forge · the domain model

Industry post-training and compression on a fully self-reliant base model, delivered inside the on-site device, producing national-expert-level decision proposals.

connect · sensing & execution

Cameras, lidar, satellite feeds and gauge data converge; uncrewed vehicles, drones, automatic devices and valves come under unified control. The machine sees, and it can reach.

decide · decision & evolution

Machine-generated response plans go to execution only after the responsible person reviews and approves; execution data flows back, effectiveness is assessed in real time, and the model keeps getting stronger.

Baseline 1 · accountability stays human — the decision always belongs to the person in charge Baseline 2 · everything traced — traceable, replayable, auditable Baseline 3 · sovereign stack — models, compute, supply chain fully self-owned Baseline 4 · stronger with use — every operation trains the system
emergency commandcritical-site securitywater & energyheavy industryindustrial parkslegal & public-sector knowledge
Deployment

Sensing and execution connect nearby.
Decisions happen on site.

A desktop box in the duty room or command center — a normal office environment is enough, no machine room. Cameras, radar and gauges plug in; uncrewed equipment comes under command.

Topobius Gradientron runtime · model pool · KV LAN < 50 ms Sensors camera · lidar · gauge Equipment uncrewed · valves · automation Command wall streaming interaction Business systems HTTP / WebSocket / SDK

The minimal ask of a site

The hub exposes standardized interfaces. Sensing and execution devices only need two things: reach the LAN, and understand the commands.

Sensing uplink
Camera / lidar / sensor / gauge data flows up as one
Command downlink
Control commands reach uncrewed equipment and automatic devices
Streaming interaction
Command wall and duty terminals at LAN RTT under 50ms
Return for evolution
Operating data and evidence flow back; models and flows keep upgrading
Open integration

Five lines of code
to your business system.

 quickstart.py
from k3box import AgentClient

client  = AgentClient("http://hub.topobius.local:8080")
session = client.session(agent="site-command", terminal="coa-01")
reply   = session.chat("Assess the situation; propose three options")
# → the machine proposes, the person decides

Python / Node.js / Java SDKs, HTTP-JSON for existing systems, WebSocket streaming output. Every instruction and every disposition is traceable and auditable.

Management console

Open a browser.
The whole box, in view.

Session monitoring
Per-agent status, token usage, latency and planning steps in real time
Model management
Switch active models, upload fine-tuned weights, tune inference parameters
Knowledge upkeep
Documents & FAQs hot-update — no restarts, no broken sessions
Alert center
Response timeouts, memory risk, link loss — email & SMS reach-out
Ops loop
OTA staged rollout & rollback, auto-inspection reports, full-chain logs, one-click backup/restore
Analytics
Session volume, hot questions, task success rate, satisfaction trends
Three routes, one ledger

Sixteen concurrent agent streams — cost and risk, side by side.

Topobius GradientronPer-terminal deploymentPublic-cloud API
Intelligence levelFrontier-classCompromisedFrontier-class
Hardware1 deviceOne per person, device churnNone
Data boundaryStays in-domainStays on terminalLeaves to cloud
Network dependenceRuns offlineRuns offlineHard dependency
Latency< 50 ms (LAN)< 50 ms (local)100–500 ms (WAN)
Ops1 device, centrally managedWalks out with the personVendor-dependent
Running cost≈ ¥50/mo electricity≈ ¥80/mo electricityPer-token, volatile
Compliance riskLowLowMedium-high (data export)
Chassis specs

A desktop body.
A machine-room watch.

CNC-milled all-aluminum unibody, at home in the duty room, command center or office. Key component figures follow the formal delivery list.

Volume11.1 L (155 × 210 × 340 mm)
Weight4.2 kg · all-aluminum unibody
ThermalActive thermal design
EnvironmentDesktop-class · no machine room
Power220 V
Operating temp0–50 °C
NetworkDual uplink (primary / backup)
ExpansionHigh-speed storage · wireless module bay
Front I/OUSB-C / USB-A
Supply chain & compliance

From the transistor to the runtime,
on a controlled chain.

Full-stack sovereign supply
Processor, chassis, inference framework, models and agent runtime — self-owned end to end
MLPS Level 2
Deployment configuration baseline for China's MLPS Level 2 provided
Personal information
All personal information processed locally, compliant with PIPL
Transport security
TLS 1.3 in transit, OAuth 2.0 on APIs
Auditable
Model behavior fine-tunable and traceable; full-chain request logs kept
Fleet deploymentFLEET DEPLOYMENT · IDENTICAL BUILD
Three-quarter viewTHREE-QUARTER VIEW · WARM DESK
High angleHIGH ANGLE · DESKTOP PRESENCE
See it on site, believe it on site

Put the next intelligent system
where your eyes are.

Book an on-site demo: bring one of your existing terminals, plug it in, and watch the first-token latency, the concurrency, the data flow.