Sentiment labels are useful for your brand when they agree with human reviewers on the messages you receive, and their mistakes fit the decision you plan to make. Test them on a held-out sample across your relevant languages, channels and contexts. Measure missed criticism and false alarms separately before letting labels change a report or a response queue.
There is no single social media sentiment analysis accuracy figure that answers this question for every brand. A tool might sort routine praise well while missing complaints about delivery. Your evaluation needs to reveal that difference.
Decide what the label will control
Write one intended use before collecting examples. For instance, you might use negative labels to prioritize a queue that a person will review. That makes two errors worth counting:
- A negative comment receives another label and never reaches that queue.
- A neutral or positive comment receives a negative label and takes review time.
Set acceptable limits for each error before seeing results. These are business choices. A monthly summary and an urgent response queue may need different limits.
Specify the target, too. Are reviewers judging the brand, a product feature or the whole message? A comment can praise a creator while criticizing the sponsored product. Scoring its overall tone may answer a different question from the one your team asked.
Sprout Social's sentiment guide discusses both aspect-level sentiment and challenges such as sarcasm and regional language. Use those challenges to define your test cases. Its overview does not establish how well a particular system handles your comments.
Build a sample from the input you will use
Start with a written collection boundary: platforms, accounts or queries, date range, languages, available context and exclusions. Save the collection settings alongside the model version and evaluation date.
Collection limits affect what you can test. For example, YouTube's commentThreads.list documentation requires a video, channel-related or thread-ID filter. Search terms narrow the returned comments; ordering and pagination affect which results you retrieve. Moderation filters require authorization, and disabled comments or insufficient permissions can prevent retrieval. A selection of accessible comments cannot establish what all customers think.
Use two separate sets:
- Routine sample. Randomly select messages from the defined collection. Include each operational language and channel. If you deliberately sample extra messages from a small language group, report that group's results separately. Do not present the unweighted combined score as everyday performance.
- Challenge sample. Deliberately gather cases involving slang, negation, comparisons, mixed-language text, sarcasm and missing context. Use this set to discover failure types. Its error rate does not estimate how often those failures occur in the routine feed.
Keep related messages from the same thread together when separating development and evaluation data. Repeated copies of a message should not become independent evidence of quality. Remove duplicate posts from a listening sample explains the collection decision to make before scoring.
Reserve an untouched evaluation set. Use separate practice messages to revise instructions or prompts. Once you change the system after inspecting evaluation errors, use fresh held-out messages for the next acceptance decision.
There is no universal sample size here. Start with a volume your reviewers can finish, publish the count for each group and expand thin groups before relying on their percentages. A handful of negative examples cannot support a stable estimate of missed criticism.
This design applies NIST's Measure guidance: document test sets and metrics, test under conditions similar to deployment, and involve relevant experts. This sampling plan is an editorial recommendation, not a platform requirement.
Create a human reference you can inspect
Have two reviewers independently label the same evaluation messages without seeing the model output. Choose reviewers who understand the original language and relevant regional usage. If the production system translates messages, evaluate that complete workflow against reviewers reading the original text.
Give reviewers the same permitted context that the system receives. Where the original post or an image is missing, record that limit. Do not silently give reviewers a full thread while testing a model that sees one sentence.
Use this starting rubric:
| Field | Rule for this evaluation |
|---|---|
| Target | Name the brand or feature being judged before assigning sentiment. |
| Positive | The message expresses approval toward that target. |
| Negative | The message expresses disapproval toward that target. |
| Neutral | The message is relevant but expresses neither approval nor disapproval. |
| Mixed | The message expresses both approval and disapproval toward the target. |
| Unclear | Available context or language knowledge does not support a judgment. |
| Out of scope | The message concerns another target or fails the collection rules. |
These are reference labels. If the tool supports only three classes, decide in advance how to handle mixed and unclear messages. Keep their counts visible. Reporting performance only on clear cases requires a separate statement of how much input those cases cover.
Record initial reviewer agreement before resolving differences. Agreement equals matching initial labels divided by all double-reviewed messages. It helps diagnose the rubric; it does not prove the reviewers are correct.
Keep a disagreement log. These rows are synthetic examples, not public posts:
| Item | First labels | Cause | Resolution for reference set |
|---|---|---|---|
| S01 | Positive / mixed | Praise for product, criticism of delivery | Mixed under the declared whole-brand target |
| S02 | Neutral / unclear | Reply refers to an unavailable image | Unclear; preserve missing-context flag |
| S03 | Positive / negative | Regional slang interpreted differently | Language-qualified reviewer adjudicates with rationale |
An adjudicator should record the reason for each final decision, preserving both initial labels. If disagreements cluster around the definition of the target, revise that definition before scoring the model. For case-level instructions, use the sarcasm and mixed-sentiment procedure.
Read the confusion matrix before the headline score
A confusion matrix counts where human reference labels and model predictions agree or differ. This hypothetical table contains 100 clear, in-scope messages. It excludes mixed and unclear cases solely to demonstrate three-class arithmetic. It reports no vendor result.
| Human reference | Model positive | Model neutral | Model negative | Total |
|---|---|---|---|---|
| Positive | 36 | 3 | 1 | 40 |
| Neutral | 4 | 32 | 4 | 40 |
| Negative | 2 | 6 | 12 | 20 |
| Total | 42 | 41 | 17 | 100 |
Using Google's definitions of accuracy, precision and recall:
- Overall accuracy is the diagonal total divided by all messages: 80/100, or 80%.
- Negative-label precision is correct negative predictions divided by all negative predictions: 12/17, about 70.6%.
- Negative-label recall is detected negative messages divided by all reference-negative messages: 12/20, or 60%.
Here, negative sentiment is the target class for the precision and recall calculations. Eight of the twenty negative messages were missed. Five of the seventeen negative predictions were false alarms. An 80% headline hides both costs.
The table also shows a reporting problem. The model predicts seventeen negative messages when reviewers identify twenty. Some missed negatives and false alarms cancel in the total. Similar sentiment totals can therefore coexist with wrong labels on individual messages.
Repeat this table by language and operational context, showing counts beside percentages. Compare tools on the same held-out messages and reference labels. Keep challenge-set results separate, and report abstentions or unclassified outputs rather than discarding them.
Make a scoped use decision
Use the results to choose what happens next:
- If reviewer disagreement remains unresolved, repair the reference set before ranking tools.
- If a language has too few examples, keep its labels provisional and review its messages manually.
- If missed criticism exceeds your limit, do not use sentiment as the sole route into a response queue.
- If performance meets your limits, run the proposed workflow with human review and record new error types.
For alerts, connect the test result to mention-alert review capacity. A usable label still needs someone to inspect the messages it selects.
Record the approved use, excluded contexts, sample counts, error limits, model version and review owner. Recheck after changes to the model, query, language mix or campaign context. Your next action is to select one decision and write its acceptable missed-message and false-alarm limits before labeling the first sample.



