Skip to content

Document Loaders

Document Loaders turn a source — a web page, a file, or the currently-open browser tab — into a plain text string. That text is meant to be wired into the Indexer node’s Input port.

The Indexer takes a plain string on its Input port and, using its text-splitter dependency, turns that string into Documents internally before embedding them into a vector store. The loaders never touch the Indexer or its dependencies — each one simply outputs text under a configurable field (default text), which you connect to the Indexer’s Input.

[ Document Loader ] --text--> [ Indexer ] --> vector store

Because a loader’s output is ordinary text, you can also route it through any other node that consumes text (an AI Agent, a summarizer, a data transform) — the Indexer is just the most common destination.

LoaderSourceRuns where
URL Document LoaderA web page fetched over HTTPWeb app or extension
File Document LoaderA file/binary payload on the input item (PDF, HTML, text)Web app or extension
Current Page Document LoaderThe text of the currently-open tabBrowser extension only

The loaders reuse the same building blocks as the standalone parse/scrape nodes, so you can mix and match:

  • The URL loader uses the same fetch + HTML-to-text conversion as Get All Text from Link.
  • The File loader uses the same PDF extraction as the PDF Parser; for a PDF already on an item you can use either node.
  • The Current Page loader uses the same page-text extraction as Get All Text.

Use a dedicated parse node when you only need the parsed text; use a Document Loader when the goal is to feed the Indexer.