Goodfire's probes run during inference with no added latency
Original titleRunning the probes during inference maintains the same throughput with no added latency.
AISummary
Goodfire reports that running its probes during model inference maintains the same throughput with no added latency. The company attributes this to infrastructure engineering, including kernel-level optimizations and a custom inference server.
Source: Goodfire · x.comPublished · added here