Dev Playground
Practise what you actually use when building with local AI
30 exercises in your browser and a capstone per track. You write code, it checks itself and, when it fails, tells you what it expected and what it got. Hints and a solution if you get stuck. No account.
LLM engineering Requests, streaming, JSON, RAG and sampling.
0/12
- Cosine similarity Semantic search compares embeddings with cosine similarity. Implement it and cover the edge cases: different lengths and zero vectors.
- Chunk text with overlap Before indexing a document for RAG you split it into chunks. A small overlap avoids cutting an idea right in half.
- Read Ollama's stream Ollama answers line by line in JSON. Join the text and compute tokens per second from its own counters.
- An OpenAI-compatible chat request Ollama, vLLM and LiteLLM accept the same format as the OpenAI API. Build the right URL and body and validate the input before sending anything.
- Prompt templates without surprises Fill {{variables}} in a prompt, fail when one is missing, and keep user text from opening or closing template markers.
- Softmax with temperature An LLM's temperature parameter divides the logits before the softmax. Implement it in a numerically stable way and see its effect.
- Read an SSE chat stream With
stream: true, the OpenAI-compatible chat API answers with Server-Sent Events. Join the text fragments and stop at[DONE]. - Retrieve the k most similar documents The retrieval step of RAG: rank documents by similarity to the query and keep the top k, with stable tie-breaking.
- Does the model fit in memory? The same sum our product pages use: weights, KV cache and runtime overhead against usable memory.
- Fit the context window A local model with an 8K context cannot take the whole history. Keep the system message and the most recent messages that fit.
- Pull the JSON out of an LLM reply You ask for JSON and the model answers with a sentence, a ```json block and sometimes stray braces. Extract the first valid object or return null.
- Top-p (nucleus) sampling Ollama's
top_pparameter drops unlikely tokens before picking one. Implement that filter.
Reliability Retries, timeouts, caches and limits.
0/6
- Debounce: do not call the model on every keystroke An autocomplete that sends a prompt per keystroke saturates the server. Wait until the user stops typing.
- An LRU cache for model answers When two users ask the same question there is no need to generate again. Keep the latest answers and drop the least used.
- Retries with exponential backoff A busy model fails now and then. Retry without hammering the server: 100, 200, 400 ms.
- Cut slow calls with AbortController A saturated model server can take minutes. Set a time limit and actually cancel the request.
- Limit concurrent requests Sending 50 prompts at once to a local model only fills its queue. Process them two at a time and keep the order.
- Rate limiting with a token bucket Your API in front of the model should accept short bursts without letting anyone hog it. A token bucket does both.
Data and text Parse, clean, validate and escape.
0/6
- Escape model output before rendering it A model can return HTML, by accident or because someone asked for it in the prompt. With
innerHTMLthat HTML runs. - Clean slugs for Spanish titles Accents, ñ and opening punctuation end up as broken URLs. Turn them into a readable slug.
- Dedupe models and group them by family A model inventory with repeated entries: keep the latest of each one and count how many there are per family.
- Validate the JSON a model returns Asking a model for JSON does not guarantee it arrives well formed: fields go missing or numbers come back as text. Check it before you use it.
- Parse a CSV with quotes Before indexing data for RAG you have to read it correctly. A
split(',')breaks on the first name that contains a comma. - Extract ISO dates and Spanish dates Before handing text to a model, pull the dates out with a regular expression and drop the ones that do not exist.
Web basics The interface around the model: accessible components, SVG and canvas.
0/6
- An accessible toggle button A button that switches on and off must say so to screen readers too, not just with colour.
- Keyboard shortcuts for a prompt Ctrl+Enter (or Cmd+Enter on a Mac) sends the prompt and Escape clears it. Every other key keeps working as usual.
- Render a model list without injecting HTML Model names come from outside. Rendering them with
textContentstops an odd name from turning into code. - A sparkline on canvas A small line that shows the latency trend. Scale the values to the canvas height and draw the polyline.
- SVG bar chart of tokens per second Draw the measured speed of three models in an accessible SVG, each bar proportional to the fastest.
- Accessible tabs with ARIA and the keyboard Three buttons and three panels are not a tab widget until the ARIA roles say so and the arrow keys work.