Web Scrape to JSON: a workflow with 4 steps, starting with Run now and ending with Show the data.
How it works
Section titled “How it works”-
Run it
When Started has no settings. Click Test Workflow on the canvas.
Make it yours — run it every morning with a Schedule trigger.
-
Fetch the page
Get HTML from Linknode reference
Get HTML from Link downloads
https://news.ycombinator.comand puts its HTML inhtml. The page doesn't need to be open.Make it yours — change URL to the page you want. Pages that need you to be signed in may return a login page instead.
-
Pull out the items
HTML Extract in list mode finds every
.titlelineand, inside each, reads the link text astitleand itshrefaslink.Make it yours — set Container Selector to one repeated item on your page and add a field per value you want (price, date, author).
-
Show the data
Display Object shows the items as a tree you can expand.
Make it yours — save them with Google Sheets or download them with Save as File.
Test it
Section titled “Test it”-
Install the recipe and open it on the canvas. To import the file yourself, use Add Workflow › Import, choose the
.awfand click Create & open. Then click Make it mine at the bottom of the canvas, tick the box and confirm, so you can edit it. -
Click Test Workflow.
-
Open Show the data in the run results: you see about 30 items, each with a
titleand alinkfrom the Hacker News front page.
Variations
Section titled “Variations”| Variation | What to change | Level |
|---|---|---|
| Scrape the page you’re on | Replace Get HTML from Link with Get All HTML, which also works on pages you’re signed in to. | Beginner |
| Keep only new items | Store the last titles with Data Store and use Remove Duplicates. | Advanced |
| Get a CSV | Add CSV Build and Clipboard Write. | Beginner |
Troubleshoot
Section titled “Troubleshoot”- The list is empty. The site renders its list with JavaScript, so the fetched HTML has no items. Open the page and use Get All HTML instead.
- Links are relative (
/item?id=…). The site writes them that way. Prefix them with the site address in an Edit Fields step. - The fetch fails in the web app. Some sites block requests from other origins; run it in the extension. See browser compatibility.









