Compare comments across languages by coding their meaning in the original language, then mapping those decisions to shared definitions. Keep the original text, a context note, translation uncertainty and reviewer disagreements together. A multilingual listening codebook should let another reviewer reconstruct a decision without treating an English translation as the only evidence.
This is a recommended workflow for comparing themes in a collected comment sample. It does not establish what everyone in a language group thinks.
Define what the comparison means
Start with a question narrow enough to code, such as "Which parts of the returns process confuse commenters?" Define the unit as one comment about one named subject. A comment can receive several topic codes, but it should count once in a comment-level total.
Keep these dimensions separate:
- Topic: what the comment discusses, such as return eligibility or refund timing.
- Sentiment target: the product, delivery service, creator or another commenter.
- Evaluation: positive, negative, mixed, no evaluation expressed or unresolved.
- Intent: requesting information, describing an experience, recommending an action or unclear.
These distinctions prevent a compliment about a creator from becoming praise for the product. Sprout Social's sentiment guide describes aspect-based analysis and the difficulty of interpreting sarcasm. Use those distinctions when designing codes; its suggested sentiment-score thresholds are not rules for comparing languages.
Define "mixed" as evidence of both positive and negative evaluation of the selected target. Define "unresolved" as insufficient evidence to choose a label. Neither belongs in "no evaluation expressed."
Record collection differences before language differences
Attach a collection note to every batch. Record the platform, selected posts, date window, collection date, queries, ordering, pagination and inaccessible material. Also record whether the batch includes replies or only top-level comments.
For example, YouTube's commentThreads.list documentation describes video and channel filters, search-term filtering, time or relevance ordering, and pagination. Its moderation-status parameter requires a properly authorized request. A video with disabled comments can return a commentsDisabled error.
Those mechanics belong in the comparison. A relevance-ordered page in one language and a longer time-ordered collection in another are different selection procedures. Record a failed retrieval as unavailable, never as an empty conversation. This YouTube example does not establish access rules for other platforms.
Review local spellings and brand nicknames with language reviewers before collection. Use the guide to tracking nicknames and transliterations when building that query list. Otherwise, the codebook may work well on a sample that missed the expressions people use.
Do not infer a commenter's country, ethnicity or first language from the comment language. Keep language and known market context in separate fields. Mark unknown context as unknown.
Build shared definitions with local notes
Give each code a stable identifier. Write its definition, inclusion rule, exclusion rule and boundary example. Translate these instructions for reviewers, then discuss whether the translated definition asks them to make the same judgment.
Keep local expressions in language-specific notes beneath the shared definition. If an expression has no useful short English equivalent, retain it and explain its meaning in a sentence.
This follows the approach recommended by van Nes and colleagues: retain the original language during analysis where possible, consider alternative wording, and return to source-language material when checking interpretations. Their paper concerns qualitative research. The codebook below adapts those recommendations for social comments.
A code definition to copy
| Field | Example definition |
|---|---|
| Code ID | RETURN_ELIGIBILITY |
| Question | Does the comment discuss whether an item qualifies for return? |
| Include | Opened items, missing packaging, excluded categories, return deadlines |
| Exclude | Refund arrival time after an accepted return |
| Multiple codes | Add REFUND_TIMING if the comment also discusses payment timing |
| Sentiment rule | Asking whether a return is allowed does not itself establish dissatisfaction |
| Context required | Identify the product or policy being discussed when available |
| Local-language note | Record local terms for return, exchange and refund separately |
| Unresolved rule | Leave the topic unresolved if the wording could mean either returning an item or returning to a store |
| Version | Record the definition version used for each coding batch |
This is an illustrative definition. Change its boundaries to fit the business question before coding starts.
Fields for each comment
Use a linked record or spreadsheet row with these fields:
| Group | Required fields |
|---|---|
| Evidence | Comment ID, source link, posted time, collected time, original text |
| Context | Parent post or reply, relevant media context, missing context |
| Language | Observed language or languages, variety if supported, reviewer competence |
| Interpretation | Literal gloss, contextual translation, alternative interpretation |
| Codes | Topic, target, evaluation, intent, codebook version |
| Uncertainty | Reason, affected decision, evidence needed to resolve it |
| Review | Reviewer A label, reviewer B label, disagreement type, final decision, reason, decision owner |
Preserve code-switching rather than forcing every comment into one language bucket. Keep machine translations in a labelled helper field. Reviewers should be able to read the original before seeing an automated sentiment suggestion.
Run bilingual review before combining results
Assign a native-language reviewer familiar with the relevant community and a second bilingual reviewer who can read the original. A native speaker may still lack the regional or subject knowledge needed for a particular comment. Record that limit.
If no qualified reviewer is available, retain the affected comments as pending. Limit the comparison rather than presenting machine-only labels as equivalent to reviewed labels.
Use this review sequence:
- Select a pilot set from each language using a documented selection rule. Include difficult cases separately, such as short replies, mixed languages and ambiguous targets.
- Have both reviewers code the original independently with the same available context. Preserve their first decisions.
- Compare disagreements by field. Separate translation disagreements from unclear definitions and missing context.
- Discuss alternative interpretations. A reviewer who can assess the original records the final decision and reason, or leaves it unresolved.
- Revise definitions and local notes. Revisit earlier comments affected by a changed rule, then try the revision on fresh comments.
The Cross-Cultural Survey Guidelines separate translation, review, adjudication, pretesting and documentation. They also recommend recording translators' questions for review. These are useful process ideas; questionnaire translation guidance does not validate a social-listening classifier.
For automated labels, use a separate sentiment-evaluation procedure. Human agreement and model accuracy answer different questions.
Preserve ambiguity in the worked example
The following comments and circumstances are hypothetical. They illustrate coding decisions, not observed customer behavior or fixed rules about a language.
| Original comment | Contextual reading | Coding decision | Review note |
|---|---|---|---|
| "Nice bag. The zip sticks." | Praise for appearance and criticism of the zip | Product target; mixed evaluation; design and zip topics | Keep both topics under one comment ID |
| "La bolsa es bonita, pero la cremallera se atasca." | The bag looks good, but its zip gets stuck | Same shared codes as the English example | Reviewer checks both clauses against the original |
| "C'est parfait." | May express approval; context could change the reading | Unresolved in this hypothetical case | Parent reply is missing; reviewers disagree about the target and tone |
Do not make the last row positive to complete the report. Record the competing readings, the missing parent reply and the decision to leave it unresolved. Do not declare it sarcastic either.
A useful disagreement note says, "A coded product praise; B could not identify the target. Parent reply unavailable. Final evaluation unresolved." A note saying "translation issue" leaves the next reviewer with no usable explanation.
Report comparable counts and unresolved cases
Report each language's collected, coded and unresolved comment counts. State the denominator beside any percentage. If you count topic mentions, explain that one comment can contribute to multiple topics.
Check agreement on the initial independent labels, before discussion changes them. Report it separately for topic, target and evaluation, with the number of double-coded comments. Keep deliberately difficult pilot cases separate from any routine-sample estimate.
Describe findings as patterns in the collected comments. A convenience sample does not justify a population claim about speakers of that language. Explain unequal access and sampling in the report using the listening coverage-gap checklist.
Before the next batch, choose one shared code, assign its language reviewers and independently code a pilot set. Resolve the definition problems before combining the language totals.



