Read Ollama's stream
Ollama answers line by line in JSON. Join the text and compute tokens per second from its own counters.
When you call /api/generate with streaming, Ollama returns one JSON line per chunk: {"response":"Hel","done":false}. The last line has done: true and the counters eval_count (tokens generated) and eval_duration (nanoseconds).
Write parseOllamaStream(text) that takes the whole stream text and returns { text, tokensPerSecond }. Ignore empty lines. Round tokens per second to one decimal.
The variable SAMPLE already holds a real example stream.
One more check uses LIVE: real output from your own Ollama if you switch on "My Ollama", or a different sample if not. That check looks at structure, not exact words.
Challenges 0/4
- Joins the text of every line
- Computes 41.3 tokens per second (124 tokens in 3 s)
- Copes with empty lines and a stream without a final line
- Works on real output from your model (My Ollama) or another sample
function parseOllamaStream(text) {
// 1. split into lines, skip empty ones
// 2. JSON.parse each line and join the `response` fields
// 3. read eval_count / eval_duration from the final line (done: true)
return { text: '', tokensPerSecond: 0 };
}
console.log(parseOllamaStream(SAMPLE));Go deeper: the related guide →
This in production, with your data? Let's talk for 15 minutes →