Skip to content

Transformers Chat

Run a small instruct/chat model entirely in your browser with transformers.js (ONNX/WASM, WebGPU where available). The runtime-agnostic sibling of Web LLM: the same chat-model dependency, a different on-device engine — a good fit for tiny, CPU-friendly chat that runs without WebGPU or an API key.

Exposes an installed transformers.js chat model (e.g. SmolLM2 360M Instruct) as a chat-model dependency. Connect it to an AI agent node and it answers on-device — messages in, text out — with no cloud call.

  • You want local chat but the device has no WebGPU (Web LLM needs it; transformers.js falls back to WASM).
  • You want the smallest possible local chat model for quick, low-stakes replies.
  • You’re already using transformers.js models for vision/text tasks and want one engine.

For larger, higher-quality local chat with WebGPU, prefer Web LLM.

Transformers Chat works with the Tools Agent. These models have no built-in tool API, so tool calling goes through the prompt, and works with any chat model:

  1. The connected tools are described in the system prompt, with a simple format for calling them.
  2. The model calls a tool by writing that call in its reply. Models often use the format they were trained on instead, so the common ones are all understood: <tool_call>{…}</tool_call>, [TOOL_CALLS] […], <function=name>{…}</function>, a fenced ```json block, or a reply that is only the JSON call.
  3. The agent runs the tool and sends the result back as a <tool_response> message; the model then answers or calls another tool.

How reliably a model calls tools depends on the model: instruct models trained for tool use do much better than very small general ones, which often just answer in text. Reasoning (<think>…</think>) is removed from the reply before it’s read.

SettingNotes
ModelAn installed transformers.js chat model. The picker marks installed models and can download others; Manage models opens Local AI.
TemperatureSampling temperature (higher = more varied; 0 = deterministic).
Max new tokensCaps the reply length (default 256).

Returns a chat-model dependency you connect to AI agent nodes.

  • “isn’t a transformers.js chat model” — the picker only lists chat models; if you selected one from another engine, use Web LLM instead.
  • Slow or short replies — these models are tiny; lower expectations, or raise Max new tokens for longer output.
  • The agent never calls a tool — the model didn’t write a valid <tool_call> block. Use an instruct model trained for tool use, give tools clear descriptions, and keep the system message short. A call to a tool that isn’t connected, or with invalid JSON, is kept as plain text instead of being run.
  • Tool calls get cut off — raise Max new tokens so the whole <tool_call> block fits in the reply.
  • First run is slow — the model loads (and downloads if needed) on first use, then stays warm for the session.