// INFERENCE CLUSTER
LIVE

iHoree LLM Inference

real-time metrics across 8 endpoints · 1h rolling history · server polls every 1s

ONLINE
0/8
CLUSTER GEN
0 tok/s
TOTAL TOKENS
0
DECODE CALLS
0

Time Series Detail

CLICK AN ENDPOINT CARD OR BUTTON TO FOCUS

Generation Throughput
llamacpp:predicted_tokens_seconds · tok/s
--
Prompt Throughput
llamacpp:prompt_tokens_seconds · tok/s
--
Cumulative Tokens
prompt_tokens_total + tokens_predicted_total
--
Requests & Slots
processing · deferred · busy_slots_per_decode
--
Copyright (c) 2026 iHoree
Made with by MOVZX