Reliability · Medium · 20 min

Limit concurrent requests

Sending 50 prompts at once to a local model only fills its queue. Process them two at a time and keep the order.

Write mapLimit(items, limit, fn). Call fn(item, index) for every item, with at most limit calls in flight at once, and return an array of results in the same order as items, even when they finish out of order.

You already have fakeGenerate(prompt, ms), which simulates a model call: it takes ms milliseconds and returns the prompt in upper case. It records the peak number of simultaneous calls in window.__peak.

Challenges 0/4

  • Keeps the order even when calls finish out of order
  • Never more than 2 calls at once with limit = 2
  • With fewer items than the limit, runs them all
  • An empty array returns []

async function mapLimit(items, limit, fn) {
  // start at most `limit` workers; each takes the next index until none are left
  return Promise.all(items.map((item, i) => fn(item, i)));
}

mapLimit(['hola', 'que', 'tal', 'modelo'], 2, (p) => fakeGenerate(p))
  .then((r) => console.log(r, 'peak:', window.__peak));
Console output appears here (console.log).

Go deeper: the Ollama reference →

This in production, with your data? Let's talk for 15 minutes →