Instagram caption scraper: collect captions by hashtag
Published
An Instagram caption scraper collects the text people wrote under their posts, one row per post, so you can read, sort and count it in a spreadsheet instead of scrolling. The fastest way to get captions in bulk is to start from a hashtag: name the tag, set how many posts and how far back, and export the rows. Each row should carry the caption itself, the hashtags pulled out of it, the post link, the date and the author, because a caption without its date and link is a quote you cannot check.
This guide covers where captions come from, what a good caption row looks like, how to run a collection from start to export, and what to do with the text afterwards. It uses Instagram Hashtag Scraper, an Apify actor that returns the recent posts under one hashtag with their captions, as the worked example, and it says when copying a handful of captions by hand is the better plan.
Why collect captions at all
Captions are where people say what a post is about in their own words. The picture tells you what something looks like; the caption tells you how its author talks about it. That makes captions useful for jobs that have nothing to do with images:
- Copywriting research. How do people in your niche open a caption, ask a question, or close with a call to action?
- Hashtag mining. Which other tags travel with the one you follow? Captions are where those tags live.
- Voice of the customer. The words buyers use for a product, a problem or a place, before a marketer rewrote them.
- Topic and sentiment work. A column of short texts is the input most text tools expect. We cover scoring it in Instagram sentiment analysis.
- Campaign records. For a branded tag, the captions are the record of what entrants actually said.
All five want the same thing: many captions, in rows, with enough context per row to trace each one back to its post.
Where an Instagram caption scraper starts
A caption belongs to a post, so a caption scraper is really a post scraper that keeps the text. There are three places it can start, and the choice decides which tool you need. Our Instagram post scraper guide compares the tools by starting point in detail.
| Starting point | Gives you captions from | Good for |
|---|---|---|
| A hashtag | Everyone who used the tag recently | Niche language, campaigns, events, hashtag mining |
| A profile | One account's own posts | Studying a competitor's or a creator's voice |
| A list of post URLs | Exactly the posts you already picked | A second pass over posts found another way |
For most caption research the hashtag is the right start, because the point is usually to hear many voices on one topic. A profile tells you how one brand writes; a tag tells you how a whole community does.
What a caption row should contain
The caption alone is rarely enough. Here is what each row from Instagram Hashtag Scraper carries, checked against its output schema today:
| Field | What it holds | Why it matters for caption work |
|---|---|---|
caption | The full caption text, or empty when the post has none | The thing you came for |
hashtags | Every hashtag found in that caption, as a list | Hashtag mining without writing a regex |
url | The post's own link | Lets you check any quote against the source |
takenAt | When the post was published | Lets you sort, filter and compare periods |
ownerUsername | The account that posted it | Lets you count distinct voices, not posts |
likeCount, commentCount | Counts at the time of the run, or empty when Instagram did not return one | A rough sense of which captions landed |
mediaType | Photo, video or carousel | Reels and photos are often captioned differently |
An empty count is not a zero. When Instagram does not return a like or comment count, the row says so with an empty value rather than writing 0, so an average you compute will not be dragged down by posts that were simply missing a number.
A post with no caption still comes back as a row, with the caption empty. That is worth knowing before you count: the share of uncaptioned posts under a tag is itself a finding, and dropping them silently would hide it.
How to scrape Instagram captions by hashtag, step by step
The run below takes a few minutes of setup the first time and very little after that.
- Pick the tag and the window. Decide which hashtag and which period you want. "Posts under #sourdough from the last 30 days" is a clear request; "everything about bread" is not.
- Get a session cookie. Instagram only serves hashtag posts to a signed-in browser, so the actor needs the
sessionidandcsrftokencookies from an account you control. The getting started page shows where to find them in your browser's developer tools. Use an account you are willing to have rate limited. - Open the actor on Apify and fill in the inputs: the hashtag (with or without the
#), the session cookie, a maximum number of posts, and the "only posts newer than" date. The date accepts a plain date like2026-09-01or a window like30 days. - Set the limit you actually want. The maximum is an exact ceiling: the run never delivers or charges for more posts than that number. Start small, say 50, to see the shape of the data before a large run.
- Run it. The run stops when it reaches the limit or the date, whichever comes first, and its status message says which.
- Export. Apify lets you download the results as CSV, Excel or JSON. For caption work, CSV or Excel is usually right; our guide to exporting Instagram posts to Excel covers the details, including how the
hashtagslist lands in a cell.
You pay only for the rows actually delivered. Duplicate posts are dropped across the whole run and are not charged, which matters on busy tags where the same post can appear on more than one feed page.
Five things to do with the captions once you have them
A spreadsheet of captions is raw material. These are the five uses that pay back the effort fastest, all doable with ordinary spreadsheet functions.
1. Build a swipe file of openings
Sort by likeCount (high to low), then read the first line of the top 30 captions. Copy the openings you like into a second sheet with the post link beside each one. Patterns appear fast: a question, a number, a confession, a "how I" line. This is research, so keep the links and write your own lines rather than reusing anyone's words.
2. Find the tags that travel together
Split the hashtags column into one tag per row and count each tag. The tags that show up beside yours most often are its neighbours, and a good shortlist for hashtag research. Ignore the very generic ones (the tags that appear under every topic) and look at the middle of the list.
3. Collect the words customers use
Search the caption column for your product category or the problem it solves. Filter to the matching rows and read them in full. The phrases people repeat are the phrases to use on your own page, because they are how buyers already describe the thing.
4. Compare two periods
Run the same tag twice with different date windows, for example the month before a launch and the month after, using "only posts newer than" and "only posts older than" together. Then compare the most common words or tags in each set. Our guide to tracking hashtag performance shows a week-by-week version.
5. Hand the column to a text tool
Most topic, keyword and sentiment tools accept a column of short texts. The caption column is that column, and the url beside it means any surprising result can be checked against the post it came from.
When copying captions by hand is enough
A scraper is the wrong tool for some jobs, and it is worth being honest about which:
- You need ten captions, once. Open the tag in the app, read the top posts and copy what you need. Setup would take longer than the job.
- You are studying one account. A profile scraper, or simply reading the profile, fits better than a hashtag run. This actor starts from a hashtag only.
- You need comments, not captions. The caption is what the author wrote. Replies under the post are comments, and this actor returns their count but not their text.
- You cannot use a session cookie. Hashtag data on Instagram needs a signed-in session. If you cannot provide one from an account you control, a scraper that depends on it will not run, and searching without an account is the realistic limit.
A note on fair use of what you collect
Captions are written by real people. Use them to learn, count and quote with attribution, not to republish someone's words as your own. Keep the post link with every caption you keep, and if you plan anything beyond research, read our overview of whether scraping Instagram is legal first.
Summary
To scrape Instagram captions in bulk, start from a hashtag, set an exact limit and a date window, and export rows that keep each caption beside its link, date, author and hashtags. Instagram Hashtag Scraper does that for one hashtag per run and bills only for the posts it delivers. For a handful of captions, the app and a copy-paste are still the quickest way.