Remove repeated collection rows by post ID, then separate unchanged reposts from new commentary. Keep quotes, replies and updates when they add meaning. Count authors separately from posts. Social listening deduplication should leave you with a record of what people said, how often material circulated, and which accounts contributed, without treating every copy as an independent opinion.
Decide what you are counting
Keep three views of the same collected material:
- Unique posts. One record per platform and post ID, including reposts. This describes collected activity.
- Authored statements. Text or other content the posting account contributed. Separate repeated statements from new experiences or changed views.
- Distinct authors. One account per platform within the declared topic and reporting period. This shows how widely participation is spread within your sample.
A quote can add a new statement while sharing an existing post. An author can contribute several useful updates without becoming several people. A repost can show circulation without revealing why its author shared it.
These are recommended analysis rules. Platforms do not define a universal unit called an independent opinion. Even distinct accounts can influence one another, belong to the same person, or publish coordinated material. Deduplication alone cannot establish statistical independence.
Keep identity and context before matching text
Preserve the original permitted export and build a separate analysis table. Keep the platform, post ID, account ID, publication time, collection time, post URL, full available content and parent-post relationship. Add a duplicate-group ID, decision reason and retained-record ID so someone can reverse each exclusion.
For example, the X data dictionary documents id, author_id, created_at and referenced_tweets. The references connect related reposts, quotes and replies. It also documents edit_history_tweet_ids for versions of an edited post. Additional fields need to be requested; an export containing only text may lose the evidence needed to distinguish these cases.
Use the platform and post ID together as the first matching key. If overlapping queries return the same ID, keep one analysis record and preserve both query labels. Differences in collection-time engagement counts do not create another post.
Keep edit versions linked. Choose a stated rule, such as the latest available version at the reporting cutoff, while retaining the permitted history for review. Do not count an edited sentence as another author.
A shared conversation ID is context, not a duplicate key. Replies within one conversation may disagree with each other.
A decision table with synthetic posts
The following hypothetical export contains eight rows about a fictional bottle lid. Times, accounts, IDs and wording are invented. Assume the reviewer can see the stated relationships. Each row is evaluated against the other rows, within one reporting day.
| Row | Record | Content or relationship | Decision for opinion analysis |
|---|---|---|---|
| A | Account A, post 101, 09:00 | "The lid leaks when I tip the bottle." | Keep the initial statement. |
| B | Account A, post 101, 09:00 | Same post collected by another query | Remove the extra collection row. Link it to A. |
| C | Account B, post 102, 09:05 | Native repost of post 101, with no added content | Record circulation. Do not assign a new opinion. |
| D | Account C, post 103, 09:10 | Copies A verbatim and credits A | Group with A as copied material. Do not infer personal experience. |
| E | Account D, post 104, 09:20 | Quotes A and adds "Mine stays dry after tightening the seal." | Keep D's contribution. Read the quote as context. |
| F | Account E, post 105, 09:30 | Replies "My bottle also drips inside my bag." | Keep the separately described experience. |
| G | Account A, post 106, 10:00 | Repeats A unchanged | Preserve the event. Exclude the repeated statement from the distinct-statement view. |
| H | Account A, post 107, 12:00 | "The replacement seal fixed it." | Keep the update, linked to A's earlier complaint. |
Here is the count produced by those rules:
| Measure | Hypothetical result | Included rows |
|---|---|---|
| Export rows | 8 | A through H |
| Unique post IDs | 7 | A, C, D, E, F, G, H |
| Retained statements after grouping copies and repeats | 4 | A, E, F, H |
| Accounts contributing retained statements | 3 | Accounts A, D, E |
The four statements include two stages of Account A's experience. Reporting four independent customers would misstate the sample. If your question concerns the latest reported state, use H for Account A at the cutoff. If it concerns the path to resolution, retain A and H as a sequence.
Match conservatively when relationships are missing
Text matching helps find candidates for review. Start with harmless formatting differences such as repeated spaces and line breaks, while retaining the original text.
Do not remove negation, emoji or punctuation before deciding whether two comments mean the same thing. "Leaks" and "Doesn't leak" must remain separate. Short generic comments such as "Great lid" are too common to prove copying on their own.
For longer near-matches, compare the added words, linked source, media and available posting context. A pasted announcement with no added position belongs in a copied-content group. The same announcement with a specific objection adds a statement worth coding. Shared links alone do not establish duplicate opinions.
Screenshots and copied captions need the same review. Matching an image does not establish that the surrounding commentary matches. When the parent post or complete content is unavailable, mark the relationship unresolved. Do not silently classify it as original or duplicate.
Keep identical wording by different accounts distinguishable in the participation view, even when you group it in the content view. Call it repeated wording unless you have evidence of copying. Similarity alone does not establish bots or coordination.
For quotes, assign sentiment to the added contribution toward a named subject. Account D's comment in the table should not inherit the complaint embedded underneath it. Use the sentiment evaluation guide to check those labels before aggregating them.
Report the collection limits beside the totals
Record your query version, platforms, time window, retrieval method, pagination completion and unavailable fields. As one platform example, X's search documentation distinguishes seven-day recent search from full-archive search with different access requirements. It lists developer-account, app and credential prerequisites. Confirm the access available to your collection method rather than assuming every tool returns the same material.
If you exclude reposts during collection, the resulting export cannot tell you how many reposts you excluded. Collect or report circulation separately when that is part of the question. Explain other missing material with a listening coverage note.
Sprout Social's query guide recommends refining search terms as conversations change. Log those changes alongside deduplication rules. Otherwise, a falling count could reflect a narrower query or a new exclusion rule. For irrelevant matches, review the brand-query exclusion procedure before changing your duplicate rules.
Deduplicated search results remain a collected sample. Pew Research Center's 2019 Twitter study used probability-based panel recruitment, account validation and weighting. It also distinguished users from their unequal posting activity. Your keyword export does not acquire that sampling design when duplicate rows disappear. Describe shares within the collected sample, not the prevalence of an opinion among customers or the public.
Before your next report, have a second reviewer inspect the proposed exclusions, especially quotes, short comments and updates that change an earlier view. Resolve disagreements in the decision log. Then publish unique-post, retained-statement and distinct-account totals with their definitions and the same reporting cutoff.



