AI Model Services
Software-hardware co-design with unified training & inference, offering LLM services, AI all-in-one appliances and token-based compute.
AI Stack All-in-One Appliance
Software-hardware synergy with unified training & inference to build AI services fast.
- 1.5TB+ VRAM
- 16-card single node, 1.5TB+ VRAM
- 700GB/s inter-card interconnect
- 1.6T bandwidth
- 1.6T inter-node communication
- Low-latency congestion-free networking
- Single node runs full DeepSeek BF16 at high concurrency
- 3000+ Tokens/s throughput
| Model | Precision | VRAM | Throughput (tokens/s) | Concurrency |
|---|---|---|---|---|
| DeepSeek-R1/V3 | BF16 | 1536 GB | 3708 | 256 |
| DeepSeek-R1/V3 | INT8 | 1536 GB | 5872 | 512 |
Note: data from lab tests; actual performance depends on delivery environment.
On-premise private deployment keeps sensitive data (government, finance) off the cloud, meeting data-sovereignty and privacy requirements.
Pre-installed optimized models and full-stack toolchain (data processing, distillation fine-tuning, agent building) enable out-of-the-box use, cutting deployment from weeks to hours.
16-card node at 16/8/4-bit supports high-concurrency full-power DeepSeek-R1/V3; BF16 8K+ tokens input latency within 50ms; in-house engine boosts throughput 50% and cuts latency 50% vs vLLM.
Pay-per-Token Elastic Compute
No self-built clusters; pay for actual usage and call LLM inference instantly.
Token-based metering; transparent and controllable cost.
Scale compute on demand to handle peaks.
Unified access to DeepSeek and mainstream LLMs.
Optimized inference pipeline for stable low latency.
Start Your AI Transformation
Whether private all-in-one or elastic token compute, we match the best AI landing path for you.



