Local AI
Local AI
Section titled “Local AI”Local AI nodes run AI models entirely in your browser — no API key, no cloud call, no per-token cost, and your data never leaves the device. They are the workflow counterpart of the Local AI management page, where you install and manage the models these nodes use.
The runtimes
Section titled “The runtimes”Every Local AI node is runtime-agnostic: you pick a model, and the model decides which on-device engine runs it. Three engines are supported, chosen automatically per model:
- WebLLM — chat / text-generation models on WebGPU (Llama, Qwen, Phi…).
- TensorFlow.js — vision and embedding models (MobileNet, COCO-SSD, Universal Sentence Encoder), WebGPU with a WebGL/WASM fallback.
- transformers.js — the widest task zoo (classification, zero-shot, detection, segmentation, depth, captioning, Q&A, summarization, translation, Whisper speech-to-text) via ONNX.
You never choose the engine directly — the model picker lists installed models across all three, and the node resolves the right runtime for whichever you select. Each picker only lists models that can do the node’s task.
Where transformers.js runs
Section titled “Where transformers.js runs”transformers.js models run on WASM (your CPU) using compact quantized files. This is deliberate: those files give wrong results on WebGPU in current browsers, and a quietly wrong answer is worse than a slower one. WebLLM chat models and TensorFlow.js models still use your GPU when available.
The nodes
Section titled “The nodes”| Node | Task | Typical model |
|---|---|---|
| Classify Image | Label an image | MobileNet, ViT |
| Detect Objects | Objects + bounding boxes | COCO-SSD, DETR |
| Caption Image | Describe an image in words | ViT-GPT2 |
| Analyze Text | Sentiment / toxicity / moderation | DistilBERT-SST2, toxic-bert |
| Classify Text | Sort text into your own labels (zero-shot) | DeBERTa zero-shot |
| Summarize | Condense long text | DistilBART CNN |
| Translate | Translate between languages | NLLB-200, opus-mt |
| Answer Question | Extract an answer from a passage | DistilBERT-SQuAD |
| Generate Text | Summarize, translate, transform | DistilBART, NLLB, T5 |
| Transcribe Audio | Speech-to-text | Whisper tiny/base |
| Run Local Model | Any task, any installed model — including image segmentation and depth maps | SegFormer, Depth Anything |
On-device agents and knowledge
Section titled “On-device agents and knowledge”Some AI agents can also run on-device. What you connect to their Model port decides the engine:
- Local Classifier → Text Classifier or Sentiment Analysis.
- Local Q&A Model → Q&A Agent, quoting answers from its Knowledge.
- Local Embeddings → Local Knowledge, for private document search.
Before you run
Section titled “Before you run”Each node needs a model installed for its task. Open the model picker in the node and download one right there, or install ahead of time from the Local AI page. If nothing is installed for the task, the picker is empty and the node explains which kind of model to add.
The first run of a model loads it into memory (and downloads it if it wasn’t installed), which can take a few seconds; subsequent runs in the same session are fast. All inference runs off the main thread in a Web Worker, so the editor stays responsive.
Image and audio inputs
Section titled “Image and audio inputs”Vision and audio nodes give you several ways to provide the media:
- From input — an expression such as
{{ $binary }}piped from an earlier node (e.g. Record Audio → Transcribe Audio, or an HTTP download → Classify Image). - URL — a link to the image or audio.
- Upload — pick a file; it’s stored inline with the node and stays on your device.
- Base64 — paste a data URL or raw base64 (the file type is detected from the data).
- Record (audio only) — record a clip from your microphone right in the settings.
Each of these nodes also has a Run preview button: it runs the task on-device against the current media and selected model and shows the result right in the settings — top labels and scores as bars, or the generated text / transcript — so you can check a model before wiring it up.
Privacy and cost
Section titled “Privacy and cost”Because everything runs locally: there is no usage cost, the workflow works offline once models are installed, and no image, text, or audio is ever sent to a server. This makes Local AI a good fit for private documents, personal media, and moderation you don’t want to route through a third party.
Related
Section titled “Related”- Local AI page — install and manage the models.
- Record Audio — capture microphone audio to feed Transcribe Audio.
- Sentiment Analysis — runs on-device with a connected Local Classifier.