Inference Clinic / sample screening

Find the expensive part of the real path.

This report shows the shape of a first-pass review using an eight-request fixture. It is a screening artifact, not a savings promise. A paid diagnostic replaces the fixture with a representative workload and a quality metric.

Workload: sample telemetry Requests: 8 GPU rate: $2.40/hour
p95 latency2,460 ms
success rate75.0%
GPU utilization49.6%
tool calls/request1.625

Signals worth investigating

High
Low GPU utilization

Average GPU utilization is 49.6%.

Next: Check batching, request-shape fragmentation, placement, and queue starvation before adding hardware.

High
High tail latency

p95 latency is 2,460 ms.

Next: Separate prefill and decode time, then inspect long-context requests and tool-call fan-out.

High
Reliability gap

Success rate is 75.0%.

Next: Partition model, retrieval, tool, and infrastructure failures so the fix is measurable.

Medium
Tool-call amplification

Average tool calls per request is 1.625 with a 75.0% success rate.

Next: Trace retries and tool-call branches before optimizing kernels.

Want this on your workload?

Start with a $1,500 diagnostic. If there is a clear lever, the fee credits toward the $7,500 five-business-day implementation sprint.

Request a diagnostic

Method note: all figures above are generated from the local sample fixture. Validate every intervention against representative traffic, a quality metric, and an agreed decision rule.