Skip to content
Agentic Workflowdocs
v0.8.2Install free

use the app

Local AI

AI models that run entirely in your browser (chat, vision, text and audio), installed once and used by Aria, agents and workflows. Nothing leaves your device.

Free · no usage cost Runs and stays on this device New in 0.8.0

In the app: sidebar › Local AI (or Settings › Local models)

Local AI is where you find, install and manage the AI models that run entirely in your browser. Once a model is installed it works offline, costs nothing per use, and no image, text or audio leaves your device. You pick a model; AWFlow picks the engine:

  • WebLLM: chat and text-generation models (Llama, Qwen, Phi and more) on WebGPU.
  • TensorFlow.js: vision, text and embedding models (MobileNet, COCO-SSD, toxicity, the Universal Sentence Encoder), WebGPU with a WebGL/WASM fallback.
  • transformers.js: the widest range of tasks (image classification, object detection, segmentation, depth, captioning, sentiment and zero-shot classification, question answering, summarization, translation, embeddings and Whisper speech-to-text) through ONNX on WASM.
The Local AI page: storage used by on-device models split by engine (WebLLM, TF.js, Transformers), a task list on the left (Chat, Embeddings, vision and audio tasks) and model cards on the right, such as Llama 3.2 1B Instruct with its size, variants and install state. The Local AI page: storage used by on-device models split by engine (WebLLM, TF.js, Transformers), a task list on the left (Chat, Embeddings, vision and audio tasks) and model cards on the right, such as Llama 3.2 1B Instruct with its size, variants and install state.
  1. Storage header: space used, by engine
  2. Tasks
  3. Search and engine filter
  4. A model card: size, device fit and install state
  1. Storage header. How much browser storage your models use, out of what’s available, and Keep models installed.
  2. Tasks. Chat, Image classification, Object detection, Speech-to-text, Embeddings… each with Show all.
  3. Search and engine filter. Narrow to WebLLM, TensorFlow.js or transformers.js.
  4. A model card. Its size, its device fit and whether it’s installed, with Install, Test and the compare checkbox.
Where How
Chat with Aria Pick an on-device model as Aria’s model, or Set default on a chat variant. In Ask mode, Aria can also classify an image, moderate text or transcribe audio with a local model and answer from the result. If no suitable model is installed, it tells you which kind to add.
An agent Choose an on-device model on the builder’s Model tab, or turn on the On-device AI skill. See Agent skills.
A workflow The Local AI nodes, Web LLM, Transformers Chat, Local Classifier, Local Q&A Model and Local Embeddings. A model installed here is used with no re-download.

Beyond the built-in list, Add custom model takes any MLC-compiled WebLLM model, a TensorFlow.js model.json, or a transformers.js model id with an ONNX build. AWFlow detects its task, and you confirm that you trust the source before it downloads. Step by step.

Ask Aria