Instagram Hashtag Scraper

Instagram sentiment analysis: score a hashtag's captions yourself

Published

Instagram sentiment analysis means taking the text people post and labelling each piece as positive, negative or neutral, so you can see how a conversation leans and how that changes over time. For a brand, a campaign or a product category, the most useful text to score is usually the captions under one hashtag. The method has three parts: collect the posts for a fixed window, clean the captions, then score them. You can score by hand in a spreadsheet, with a free open-source library in Python, or by paying a listening tool that does it for you. This guide walks through all three and says which one fits which job.

What you can and cannot learn from captions

Before scoring anything, be clear about what the text is.

So the honest question captions answer is "how do people who chose to post under this tag talk about it". That is useful for a campaign, a launch or a category. It is not a customer satisfaction survey, and it is not the reaction in the comments.

Step 1: collect the posts for a fixed window

Sentiment only means something against a window you can name. "Last week" is useless a month later, so write the dates down.

  1. Pick one hashtag per analysis. Two tags are two datasets, joined later on the post's shortcode so a post using both counts once.
  2. Fix the dates, start and end, and the time zone.
  3. Cap the rows. Decide the most posts you will score. For a hand-scored sheet, a hundred is already a long afternoon.
  4. Keep these columns at minimum: the post link, the date it was published, the caption, the account that posted it, and the tags in the caption.

The manual way is to open the hashtag on Instagram, scroll the recent posts, and paste each caption and link into a sheet. It works for a few dozen posts. It gets slow past that, and the feed does not let you jump to a date, so a window from three weeks ago means a lot of scrolling. Our guide on searching Instagram posts by date covers the tricks that help.

The automatic way is a tool that reads the hashtag feed and writes rows. The Instagram Hashtag Scraper on Apify takes a hashtag, an onlyPostsNewerThan date (and onlyPostsOlderThan for the far end of the window) and a resultsLimit. It stops the run when the feed reaches your start date and writes one row per post with caption, hashtags, takenAt, ownerUsername, likeCount, commentCount and url. Duplicates are dropped within a run, and a post with no caption comes through with caption set to null rather than an empty string. It costs $0.0005 per delivered post, so a 1,000-post window is $0.50. It needs a session cookie from an Instagram account you control, and the getting started page walks through that.

It does not collect comments, and it does not score anything. You get the rows and do the scoring below. The limits page lists everything else it does not do.

Export the dataset as CSV or Excel. The Excel export guide shows how each format opens.

Step 2: clean the captions before you score

Scoring raw captions is how a tag full of giveaways reads as "very positive". Four passes fix most of it:

  1. Remove duplicates by shortcode, if you collected more than once.
  2. Drop your own account and known partners. Filter ownerUsername against a list you keep beside the sheet.
  3. Drop tag-only captions. If a caption has no words left after removing every #tag, it has nothing to score. Keep a count of how many you dropped, so the final number says what it covered.
  4. Strip the tags but keep the emoji. Tags like #love and #happy are attached to everything and inflate positive scores. Emoji, on the other hand, are real signal.

Language matters too. Most free scoring methods are built for English. If the tag is multilingual, split the rows by language first and score only the ones your method reads.

Step 3, option A: score by hand in a spreadsheet

For up to a couple of hundred captions, hand scoring is the most accurate option you have, because you read sarcasm, context and emoji the way the poster meant them.

  1. Add a column called sentiment with three allowed values: positive, neutral, negative. Use a data validation list so nobody types "pos".
  2. Add a topic column too. A negative caption about shipping and a negative caption about taste are different problems.
  3. Have two people score the same twenty rows first, then compare. Where you disagree, write a one-line rule ("a complaint followed by a joke is negative") and carry on.
  4. Count each value with COUNTIF, and divide by the rows you scored, never by the rows you collected.
ApproachGood forWeak at
Hand scoring in a sheetUp to a few hundred posts, sarcasm, mixed feelingsSpeed, repeating it every week
VADER in PythonThousands of English captions, weekly repeatsSarcasm, domain words, other languages
A paid listening toolMany sources at once, alerts, client reportsCost, seeing exactly how a score was made

Step 3, option B: score automatically with VADER in Python

VADER is a free sentiment library whose own README, as of October 2026, describes it as "a lexicon and rule-based sentiment analysis tool that is specifically attuned to sentiments expressed in social media". It is released under the MIT License, handles emoji, and needs no training data and no account.

Install it with pip install vaderSentiment pandas, then score the exported CSV:

import re
import pandas as pd
from vaderSentiment.vaderSentiment import SentimentIntensityAnalyzer

posts = pd.read_csv("hashtag-posts.csv")
posts = posts.drop_duplicates(subset="shortcode")
posts = posts[~posts["ownerUsername"].isin(["yourbrand", "yourpartner"])]

def clean(caption):
    if not isinstance(caption, str):
        return ""
    return re.sub(r"#\w+", "", caption).strip()

posts["text"] = posts["caption"].apply(clean)
posts = posts[posts["text"] != ""]

analyzer = SentimentIntensityAnalyzer()
posts["compound"] = posts["text"].apply(lambda t: analyzer.polarity_scores(t)["compound"])

def label(score):
    if score >= 0.05:
        return "positive"
    if score <= -0.05:
        return "negative"
    return "neutral"

posts["sentiment"] = posts["compound"].apply(label)
print(posts["sentiment"].value_counts(normalize=True))
posts.to_csv("hashtag-posts-scored.csv", index=False)

The compound score runs from -1 to +1. The cut-offs in the code are the ones VADER's README calls typical: 0.05 and above is positive, -0.05 and below is negative, and everything between is neutral.

Read twenty scored rows by hand before you trust the totals. VADER does not know your product or your niche's slang, and a caption that is mostly a recipe or a product spec scores close to neutral whatever the poster felt. Its lexicon is a plain tab-delimited text file of words and ratings, so a word that keeps misfiring in your niche can be given your own rating.

Tracking sentiment week by week

A single score is a snapshot. The useful number is the change.

The hashtag performance guide covers the weekly sheet in more detail, and Instagram social listening covers choosing which tags to follow in the first place.

When the manual way is enough

You do not need any tool if:

In those cases, paste the captions into a sheet and hand score them. It will be more accurate than any automatic score.

When a paid listening tool is the better buy

A listening suite scores sentiment for you, across many sources, with alerts when the tone shifts. Pick one when:

We compared the main options in Brand24 alternatives and Brandwatch alternatives. If your question is one hashtag on Instagram and you are comfortable with a spreadsheet or a Python script, collecting the rows and scoring them yourself is the lighter route, and you can see exactly why every post got the label it did.

All articles