AI chat comparison

AI chat comparison

How quickly does which AI respond? This demo realistically simulates how Offline, Online, and Turbo modes differ. Click on one of the example questions and observe all three modes in parallel.

Note: the responses come from a simulation, not from a real model. Response content, first token time, and streaming speed are realistically modeled.

Advanced note: the investment slider further down only affects the offline mode and shows how response time and context window change depending on hardware class from mini-PC to professional rack.

Mini-PC with GPU (24 GB VRAM)
More investment in offline hardware means faster responses and more context. The investment only affects the offline mode; cloud modes are hardware-independent.
Offline ready
TBS-Mistral Small 3.2 24B Your own server
Select a question above
First token
Total
Tokens/s
Costs
Online ready
GPT-Class Cloud Model US Cloud
Select a question above
First token
Total
Tokens/s
Costs
Turbo ready
gpt-oss-120b on Wafer Scale Special Cloud
Select a question above
First token
Total
Tokens/s
Costs

Last requests in this session

ID Timestamp Question Model Mode Tokens on Tokens off Tokens/s First token Total Costs
No request yet in this session. Start with an example question above.