Runs entirely in your browser: the model is a few MB, runs on your own CPU in this tab, and nothing you type is sent anywhere. Why Corrigan?
Settings
Speed debrief
for the email above, on this machine: 4M model, 1 threadTiming
Time to first token runs from pressing the button until the first output token exists (request parsing + tokenising + prefill + sampling). Decode speed is generated tokens divided by the time after prefill. Prefill reads the whole prompt in one go. Per-token latency is the time between consecutive tokens. The browser rounds its timers to about 5 µs when the page is cross-origin isolated (and much more when it is not), so single tokens are noisy while the medians are fair.
Where each token's time goes
Tick profile above and write an email to see where each token's time goes.
History
Every email you write is logged here with its timings.
Want a number to compare with other machines? Run the benchmark and see the leaderboard.