Accuracy & limits
What is real here?
Check which tools use real components, which are simulations, what leaves the browser, and what each view deliberately leaves out.
Understanding the visualisers
What is AI Signal Studio?
A collection of independent browser visualisers for seeing how language-model systems work. Each tool isolates one mechanism, lets you manipulate it, and states what the view leaves out.
Is an AI thinking like a person?
This site does not make that claim. A language model predicts tokens from learned patterns. Words such as “think” or “understand” can be convenient interface labels, but they should not be mistaken for evidence of human-like experience or judgement.
Are all the visualisers using a real language model?
No. The token visualiser uses a real GPT-style tokenizer. The next-token visualiser starts with a deliberately tiny statistical model so every probability is visible. The local-model mode uses WebLLM. The neural-network and RAG tools are transparent simulations and say so on the page.
Why are some diagrams much smaller than a real system?
A literal diagram of a modern model would contain far too many values to read. Each visualiser preserves one mechanism at a time, such as token boundaries, weighted connections, or retrieval, then discloses the missing complexity.
Local model and privacy
Does my prompt leave this device?
The ordinary simulations run in this page. In the local-model visualiser, the WebLLM runtime performs inference on your device; the prompt is not sent to an AI API by this app. Normal web requests still occur to load the site, the WebLLM module, and model files from their hosts.
Why must the local model download files?
The browser needs the model weights and runtime before it can calculate an answer locally. The first load can be substantial and may take time. Browser caching can make later loads faster, but storage and cache behaviour depend on the browser.
Which browsers can run the local model?
It needs WebGPU, a compatible graphics adapter, a secure page, and enough device memory. Support varies by browser, operating system, and hardware, so the local-model visualiser tests for an adapter instead of trusting the browser name. See the current WebGPU documentation.
Does the whole site work offline?
Offline use is not guaranteed. Previously cached assets may remain available, but this project does not currently install an offline service worker. The local model also needs its runtime and model files to have been downloaded first.
Why can the local answer be odd, slow, or empty?
Small local models trade capability for a size that can run in a browser. Hardware limits, model loading, sampling settings, prompt wording, and the maximum-token limit can all affect the result. An empty or failed answer is shown as an error with a recovery path.
Limits and trust
Does RAG make an answer true?
No. Production RAG commonly searches vector embeddings or combines vector and keyword search. This site's RAG walkthrough only uses visible word matching so its selection can be explained. Either method can retrieve the wrong passage, and a model can ignore or misread a good one.
Is a percentage the model’s confidence that an answer is correct?
No. The percentages in the tiny next-token visualiser are probabilities for possible next tokens inside its nine-sentence toy world. Token probability is not a calibrated probability that a statement is true.
Where can I learn more about the browser model?
The project uses WebLLM, an open-source in-browser inference engine. Its documentation and model list are the source of truth for supported model records.