Skip to content

URL Document Loader

Load a web page as text to feed into the Indexer.

Fetches the given URL and converts its HTML response into clean plain text, then outputs that text under a configurable field (default text). Connect that output to the Indexer node’s Input port to index a web page.

Use it to index the content of a public web page into a knowledge base for RAG or Q&A, without manually copying the page text.

SettingNotes
URLThe page to fetch, e.g. https://example.com.
Output field nameField that holds the extracted text. Defaults to text.
TimeoutOptional request timeout in milliseconds.
HeadersOptional request headers (toggle Send Headers).
Include input fieldsWhen enabled, the incoming item’s fields are merged into the output.

An item with the extracted page text under the chosen output field (default text).

  • Runs in the web app or the extension. Pages that block cross-origin requests may not be reachable from the web app — in that case use the Current Page Document Loader inside the extension instead.
  • Uses the same fetch + HTML-to-text conversion as Get All Text from Link.
  • A HTTP error! status: … means the server returned a non-2xx response.
  • An Invalid URL error means the URL is malformed.
  • A Request timeout means the page took longer than the configured timeout.