Instagram Hashtag Scraper

Instagram scraper for n8n: pull hashtag posts into a workflow

Published

To use an Instagram scraper in n8n, install the Apify node, add a Run Actor step that starts a hashtag scraper with a hashtag and a date window, then add a Get Dataset Items step that reads the finished run's rows. From there the posts flow into whatever n8n connects to: a Google Sheet, a database, a Slack message, an AI step. n8n has no Instagram scraping built in, so the scraper is a separate piece you call from the workflow.

This guide builds that workflow with Instagram Hashtag Scraper, an Apify actor that returns recent posts for one hashtag. It covers the node setup, the two steps that matter, the input to paste, the fields that come back, and the two mistakes that make an automated scraper run quietly wrong.

What you are building

A workflow with four parts:

  1. Something that starts it: a schedule, a manual click, or a webhook.
  2. Run Actor: starts the scraper for one hashtag and waits for it to finish.
  3. Get Dataset Items: reads the rows that run produced.
  4. Whatever you do with rows: append to a sheet, filter, notify, summarise.

Only steps 2 and 3 are specific to scraping. Everything around them is ordinary n8n.

Step by step

As of October 2026, Apify's documentation describes the integration as follows.

  1. Install the Apify node. On a self-hosted n8n, open Settings > Community Nodes, choose Install and enter the package name @apify/n8n-nodes-apify. On n8n Cloud, search for Apify in the nodes panel and click Install node. If it does not appear on Cloud, Apify's page says to check that the visibility of verified community nodes is switched on in the Cloud Admin Panel.
  2. Create the credential. Add an Apify API credential and paste your Apify API token. Apify's documentation says the API key works on both self-hosted and cloud instances, while OAuth2 is cloud only.
  3. Get an Instagram session cookie. Instagram serves hashtag data only to a signed-in session, so the scraper needs the sessionid and csrftoken values from a browser signed in to an account. Use a separate account you are willing to lose, never your personal one. Getting started shows where the two values live.
  4. Add a Run Actor step. Add the Apify node, choose the Actors resource and the Run Actor operation, and pick Instagram Hashtag Scraper. Paste the input below into Custom input, and switch Wait for finish on so the next node does not start before the rows exist.
  5. Add a Get Dataset Items step. A second Apify node, resource Datasets, operation Get Items. Set Dataset ID to the defaultDatasetId value that the Run Actor step returned. Apify's page gives this exact wiring.
  6. Add the node that uses the rows. A Google Sheets append, a database insert, a filter followed by a Slack message. Each post arrives as one n8n item.

The input to paste

In Custom input on the Run Actor step:

{
  "hashtag": "coffee",
  "resultsLimit": 50,
  "onlyPostsNewerThan": "1 day",
  "sessionCookie": "sessionid=...; csrftoken=..."
}

Three of those fields do the real work:

FieldWhat it doesWhy an automated run needs it
resultsLimitAn exact ceiling on delivered postsA busy tag cannot make one run grow without bound
onlyPostsNewerThanStops at a date: 2026-10-01 or a window like 7 daysThe run reads only the new stretch, so each run is short
onlyPostsOlderThanSkips posts newer than a dateTogether with the field above, it reads one fixed window

The date field is a stopping rule, not a filter applied afterwards. Asking for one day means the run stops once the feed reaches yesterday.

What comes back

Each row is one post, with these fields among others: the post url and shortcode, takenAt (the time it was posted), caption, the list of hashtags, likeCount, commentCount, ownerUsername and mediaType. In n8n those become columns you can map directly. The shortcode is the one to keep for the next section.

Two mistakes that make an automated scraper quietly wrong

Overlapping windows write the same post twice. If the schedule runs daily and the window is 2 days, yesterday's posts arrive in two consecutive runs. A little overlap is good, because a run that starts late would otherwise leave a gap. But the workflow has to cope with it. Before the append step, check the shortcode against what the sheet already holds, and skip any that match. In n8n that is a lookup step plus a filter, or an upsert if your database supports one.

A failed run must not look like a quiet day. A hashtag with no new posts and a scraper that lost its login both produce an empty result if you do not look closer. Instagram sessions expire, so this will happen eventually. This actor ends the run as failed when it cannot read the feed, rather than finishing with zero rows. Give the workflow an error path: Apify's page lists an Actor Run Finished trigger, and n8n's own error handling lets you route a failed node to a message. Send the failure somewhere a person reads. Otherwise a month of silence will look like a calm hashtag.

Running it on a schedule

There are two clean ways, and neither needs the other.

Whichever you pick, match the date window to the interval, as above.

What it costs

The actor charges $0.0005 per delivered post, with no fee per run and no fee per hashtag. Rows that are duplicates, malformed or past your limit are not charged. A run that finds nothing new delivers nothing and charges nothing. A daily run on a tag that gets forty new posts costs about two cents.

Two other bills sit beside it. Apify's own platform plan is a separate layer, covered in what an Instagram hashtag job costs on Apify. And n8n meters its own use. As of October 2026 its pricing page says plans are based on monthly workflow executions "regardless of complexity", with the Starter plan listed at 20€ a month billed annually for 2.5K executions. One scheduled workflow that runs daily uses about thirty executions a month, so the n8n side is usually the smaller worry. Self-hosting is the other option the n8n site offers.

When a simpler setup is enough

You do not need n8n for this in several cases.

n8n earns its place when the posts need to go through steps: filter by caption, enrich, score, route to different places, or hand to an AI step. Know the limits of the scraper too. It reads one hashtag per run and nothing else: no profiles, no followers, no other networks. Watching ten tags means ten Run Actor steps, which in n8n is one loop over a list of hashtags.

Ideas for what to do with the rows

The short version

Install the Apify node, put a Run Actor step and a Get Dataset Items step in a workflow, give the actor a hashtag, a post limit and a date window, and de-duplicate on shortcode before you store anything. Add an error path so a lost session sends a message. Check the free hashtag checker first to see whether a tag is busy enough to be worth a schedule.

All articles