LLM engineeringCapstone · Hard · 40 min

Capstone: ask Ollama for JSON without it breaking

Puts the track together: a streaming POST to Ollama, retries with growing waits, reading the NDJSON chunk by chunk and extracting the JSON safely.

Combines the 12 exercises of the LLM engineering track. It unlocks when you finish them, but you can try it now.

Write askJson(prompt, options), the function you would put between your app and a local model. options has fetchImpl (the fetch to use), model, retries and baseMs.

1. POST to http://localhost:11434/api/generate with the JSON body { model, prompt, stream: true }.

2. If the response is not ok (say a 503 because the model is still loading) or fetchImpl throws, wait sleep(baseMs * 2 ** attempt) and try again, at most retries times. When they run out, throw an error that includes the status, for example HTTP 503.

3. Read res.body with getReader() and TextDecoder: it arrives in small chunks and a line (or an accented letter) can be split across two. Join the response fields of each JSON line.

4. Return the first JSON object in the text (the model may write something before it) or throw new Error('no JSON').

mockOllama({ failures, parts, chunk }) imitates Ollama and sleep(ms) records every wait. One check uses liveFetch, which returns real output from your own Ollama if you switch on "My Ollama", or a different sample if not.

Challenges 0/7

  • Returns the JSON object from a streamed reply
  • POSTs to /api/generate with model, prompt and stream: true
  • Retries a 503 waiting 100 then 200 ms
  • Gives up after the retries with an error that says HTTP 503
  • Copes with lines cut byte by byte and text before the JSON
  • Throws 'no JSON' when the model returns no object
  • Works on the real stream from your model (My Ollama) or another sample

const OLLAMA = 'http://localhost:11434';

async function askJson(prompt, { fetchImpl = fetch, model = 'llama3.1:8b', retries = 2, baseMs = 200 } = {}) {
  // 1. POST OLLAMA + '/api/generate' with the JSON body { model, prompt, stream: true }
  // 2. a non-ok response or a thrown error: sleep(baseMs * 2 ** attempt) and try again, at most `retries` times
  // 3. read res.body chunk by chunk (a line can be split in two) and join the `response` fields
  // 4. return the first JSON object in that text, or throw new Error('no JSON')
  return null;
}

askJson('Resume el ticket', { fetchImpl: mockOllama({ failures: 1 }) }).then(console.log, (e) => console.error(e.message));
Console output appears here (console.log).

Go deeper: the related guide →

This in production, with your data? Let's talk for 15 minutes →