Skip to content

Local AI

Local AI nodes run AI models entirely in your browser — no API key, no cloud call, no per-token cost, and your data never leaves the device. They are the workflow counterpart of the Local AI management page, where you install and manage the models these nodes use.

Every Local AI node is runtime-agnostic: you pick a model, and the model decides which on-device engine runs it. Three engines are supported, chosen automatically per model:

  • WebLLM — chat / text-generation models on WebGPU (Llama, Qwen, Phi…).
  • TensorFlow.js — vision and embedding models (MobileNet, COCO-SSD, Universal Sentence Encoder), WebGPU with a WebGL/WASM fallback.
  • transformers.js — the widest task zoo (classification, zero-shot, detection, segmentation, depth, captioning, Q&A, summarization, translation, Whisper speech-to-text) via ONNX.

You never choose the engine directly — the model picker lists installed models across all three, and the node resolves the right runtime for whichever you select. Each picker only lists models that can do the node’s task.

transformers.js models run on WASM (your CPU) using compact quantized files. This is deliberate: those files give wrong results on WebGPU in current browsers, and a quietly wrong answer is worse than a slower one. WebLLM chat models and TensorFlow.js models still use your GPU when available.

NodeTaskTypical model
Classify ImageLabel an imageMobileNet, ViT
Detect ObjectsObjects + bounding boxesCOCO-SSD, DETR
Caption ImageDescribe an image in wordsViT-GPT2
Analyze TextSentiment / toxicity / moderationDistilBERT-SST2, toxic-bert
Classify TextSort text into your own labels (zero-shot)DeBERTa zero-shot
SummarizeCondense long textDistilBART CNN
TranslateTranslate between languagesNLLB-200, opus-mt
Answer QuestionExtract an answer from a passageDistilBERT-SQuAD
Generate TextSummarize, translate, transformDistilBART, NLLB, T5
Transcribe AudioSpeech-to-textWhisper tiny/base
Run Local ModelAny task, any installed model — including image segmentation and depth mapsSegFormer, Depth Anything

Some AI agents can also run on-device. What you connect to their Model port decides the engine:

Each node needs a model installed for its task. Open the model picker in the node and download one right there, or install ahead of time from the Local AI page. If nothing is installed for the task, the picker is empty and the node explains which kind of model to add.

The first run of a model loads it into memory (and downloads it if it wasn’t installed), which can take a few seconds; subsequent runs in the same session are fast. All inference runs off the main thread in a Web Worker, so the editor stays responsive.

Vision and audio nodes give you several ways to provide the media:

  • From input — an expression such as {{ $binary }} piped from an earlier node (e.g. Record Audio → Transcribe Audio, or an HTTP download → Classify Image).
  • URL — a link to the image or audio.
  • Upload — pick a file; it’s stored inline with the node and stays on your device.
  • Base64 — paste a data URL or raw base64 (the file type is detected from the data).
  • Record (audio only) — record a clip from your microphone right in the settings.

Each of these nodes also has a Run preview button: it runs the task on-device against the current media and selected model and shows the result right in the settings — top labels and scores as bars, or the generated text / transcript — so you can check a model before wiring it up.

Because everything runs locally: there is no usage cost, the workflow works offline once models are installed, and no image, text, or audio is ever sent to a server. This makes Local AI a good fit for private documents, personal media, and moderation you don’t want to route through a third party.