{"work":{"id":"9eaaaac1-0a96-4f5f-9b13-30c46e9e1346","openalex_id":"https://openalex.org/W2936695845","doi":"10.48550/arxiv.1904.09675","arxiv_id":"1904.09675","raw_key":null,"title":"BERTScore: Evaluating Text Generation with BERT","authors":null,"authors_text":"Tianyi Zhang, Varsha Kishore, Felix Wu, Kilian Q. Weinberger, Yoav Artzi","year":2019,"venue":"cs.CL","abstract":"We propose BERTScore, an automatic evaluation metric for text generation. Analogously to common metrics, BERTScore computes a similarity score for each token in the candidate sentence with each token in the reference sentence. However, instead of exact matches, we compute token similarity using contextual embeddings. We evaluate using the outputs of 363 machine translation and image captioning systems. BERTScore correlates better with human judgments and provides stronger model selection performance than existing metrics. Finally, we use an adversarial paraphrase detection task to show that BERTScore is more robust to challenging examples when compared to existing metrics.","external_url":"https://arxiv.org/abs/1904.09675","cited_by_count":2045,"metadata_source":"pith","metadata_fetched_at":"2026-08-05T02:28:24.338817+00:00","pith_arxiv_id":"1904.09675","created_at":"2026-05-09T06:50:40.894363+00:00","updated_at":"2026-08-05T02:28:24.338817+00:00","title_quality_ok":true,"display_title":"BERTScore: Evaluating Text Generation with BERT","render_title":"BERTScore: Evaluating Text Generation with BERT"},"hub":{"state":{"work_id":"9eaaaac1-0a96-4f5f-9b13-30c46e9e1346","tier":"super_hub","tier_reason":"100+ Pith inbound or 10,000+ external citations","pith_inbound_count":201,"external_cited_by_count":2045,"distinct_field_count":20,"first_pith_cited_at":"2019-08-27T08:50:17+00:00","last_pith_cited_at":"2026-07-09T09:03:42+00:00","author_build_status":"needed","summary_status":"needed","contexts_status":"needed","graph_status":"needed","ask_index_status":"needed","reader_status":"not_needed","recognition_status":"not_needed","updated_at":"2026-08-23T04:19:26.301876+00:00","tier_text":"super_hub"},"tier":"super_hub","role_counts":[{"context_role":"background","n":8},{"context_role":"method","n":7},{"context_role":"baseline","n":1},{"context_role":"other","n":1}],"polarity_counts":[{"context_polarity":"background","n":7},{"context_polarity":"use_method","n":7},{"context_polarity":"unclear","n":2},{"context_polarity":"baseline","n":1}],"runs":{"ask_index":{"job_type":"ask_index","status":"succeeded","result":{"title":"BERTScore: Evaluating Text Generation with BERT","claims":[{"claim_text":"We propose BERTScore, an automatic evaluation metric for text generation. Analogously to common metrics, BERTScore computes a similarity score for each token in the candidate sentence with each token in the reference sentence. However, instead of exact matches, we compute token similarity using contextual embeddings. We evaluate using the outputs of 363 machine translation and image captioning systems. BERTScore correlates better with human judgments and provides stronger model selection performance than existing metrics. Finally, we use an adversarial paraphrase detection task to show that BE","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"BERTScorecomputes semantic similarity between generated and reference ADRs based on contextualized embeddings from pretrained BERT models. Unlike surface-level metrics, BERTScore captures meaning and conceptual alignment even when wording differs. This makes it particularly suitable for evaluating architec- tural documentation where paraphrasing is common. [20]. BLEUmeasures n-gram precision to capture surface-level over- lap between generated and reference texts. Originally developed for machin","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"What remains inadequately explored is not merely the retrieval of more accurate information, but rather the ability to perform abstract synthesis based on retrieved evidence under controlled evaluation settings. This mismatch also poses an evaluation challenge for abstract QA. Conventional lexical and semantic metrics, such as ROUGE [17] and BERTScore [18], are too coarse-grained for long-form abstract answers, whose wording may vary substantially while differing in topical coverage. QA-based an","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The average rank improvement across all queries where adversarial modification is applied. This metric captures the magnitude of rank shifts introduced by adversarial edits. 5.4.2 Content Fidelity Metrics.Consistent with [ 56], to evaluate how well adversarial documents preserve the semantic and structural integrity of the original payload, we employ the following metrics: Semantic Similarity (SS):The average BERTScore F1 [ 58] between the original and adversarial documents, which measures token","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"For motion, we follow prior work [30,97] and report joint accuracy (MPJPE, PA-MPJPE), mesh accuracy (MPJVE), and trajectory error (MTE) for long-horizon prediction, all in millimeters, and joint rotation error (MPJRE) in degrees. For text, we adopt the linguistic metrics from MotionGPT3 [121] and report BLEU [65], ROUGE-L [53], CIDEr [81], and BERT-Score [114]. For 3D layout, we report 3D bounding box IoU (3D- IoU), precision and recall at IoU threshold 0.5 (P@0.5, R@0.5), and identity- level pr","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"We evaluate two aspects:spatial overlap, with Jaccard at K (J@K) and soft Recall (SR@K, the fraction of queries whose best Jaccard exceeds 0.8), and ranking quality, with hard Recall (R@K) and Mean Reciprocal Rank (MRR). Task 3: Trajectory Captioning.Given a trajectory, a model must generate a factual natural language caption. We evaluatelanguage qualitywith BERTScore F1 ( BS-F1) [31], ROUGE-L (R-L) [15], and METEOR [2], andspatial groundingis measured by POI Recall ( POI-R), the proportion of g","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Table 1 | Statistics summary of datasets by sampling bias. Detailed investigations of the ConcreteBatch are performed regarding sampling bias, hardness, and intra-modal similari- ties. CLIPScores are measured with a PE-Core-L-14-336 [46] pretrained on the MetaCLIP-5.4B dataset [47] (total 58B sam- ples seen), DINOv2 [48] is employed for DINOScores, and BERTScores [49] is calculated using the default RoBERTa- large [50] with baseline rescaling. For notational convenience, let Dℎ𝑐, D𝑙𝑐 and D𝑤𝑜 den","claim_type":"method","confidence":0.8,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks BERTScore: Evaluating Text Generation with BERT because it crossed a citation-hub threshold. Current citing contexts most often use it as method evidence (7 contexts).","role_counts":[{"n":7,"context_role":"method"},{"n":6,"context_role":"background"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-05-20T06:41:48.719506+00:00"},"author_expand":{"job_type":"author_expand","status":"succeeded","result":{"authors_linked":[{"id":"42670f71-d0ea-4278-9b4e-2046cc00cf9e","orcid":null,"display_name":"Tianyi Zhang"},{"id":"b627ff25-180a-48fd-bc53-854cd6654549","orcid":null,"display_name":"Varsha Kishore"},{"id":"31d5a154-a832-4674-afc0-4a3d03df02f4","orcid":null,"display_name":"Felix Wu"},{"id":"a6bd0858-91d6-47df-8ce5-e08b05941e15","orcid":null,"display_name":"Kilian Q. Weinberger"},{"id":"1a505e27-523b-4a50-bb24-88475db88b78","orcid":null,"display_name":"Yoav Artzi"}]},"error":null,"updated_at":"2026-05-20T06:41:49.103367+00:00"},"context_extract":{"job_type":"context_extract","status":"succeeded","result":{"enqueued_papers":25},"error":null,"updated_at":"2026-05-14T08:57:52.324719+00:00"},"graph_features":{"job_type":"graph_features","status":"succeeded","result":{"co_cited":[{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":12},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":11},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":10},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":9},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":8},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":7},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":7},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":6},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":6},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":6},{"title":"Bleurt: Learning robust metrics for text gener- ation","work_id":"55ab59c9-b06b-48c6-829e-64343cc28679","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"LoRA: Low-Rank Adaptation of Large Language Models","work_id":"0426219a-789e-4964-adc8-a04538510818","shared_citers":5},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":5},{"title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks","work_id":"27adfcc9-2a67-43d6-a844-78309012411f","shared_citers":5},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":4},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":4},{"title":"Gemma: Open Models Based on Gemini Research and Technology","work_id":"a9ea2870-df28-40b8-a9e0-a7e9a116f793","shared_citers":4},{"title":"G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment","work_id":"6bc7f654-ec90-4d95-b023-416717a28f45","shared_citers":4},{"title":"Gptscore: Evaluate as you desire.arXiv preprint arXiv:2302.04166","work_id":"13d9a137-d114-4c30-aa62-3b8d25d9eb7c","shared_citers":4},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":4},{"title":"Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena","work_id":"d0c30cd7-81e1-4159-a87f-f6adca77ff08","shared_citers":4},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":4},{"title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","work_id":"bab684a8-d933-426c-a19e-2c855a0d1f59","shared_citers":4}],"time_series":[{"n":1,"year":2019},{"n":1,"year":2023},{"n":1,"year":2024},{"n":66,"year":2026}],"dependency_candidates":[]},"error":null,"updated_at":"2026-05-14T09:08:01.855528+00:00"},"identity_refresh":{"job_type":"identity_refresh","status":"succeeded","result":{"items":[{"title":"Qwen3 Technical Report","outcome":"unchanged","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","resolver":"local_arxiv","confidence":0.98,"old_work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e"}],"counts":{"fixed":0,"merged":0,"unchanged":1,"quarantined":0,"needs_external_resolution":0},"errors":[],"attempted":1},"error":null,"updated_at":"2026-05-14T08:57:56.607611+00:00"},"role_polarity":{"job_type":"role_polarity","status":"succeeded","result":{"title":"BERTScore: Evaluating Text Generation with BERT","claims":[{"claim_text":"We propose BERTScore, an automatic evaluation metric for text generation. Analogously to common metrics, BERTScore computes a similarity score for each token in the candidate sentence with each token in the reference sentence. However, instead of exact matches, we compute token similarity using contextual embeddings. We evaluate using the outputs of 363 machine translation and image captioning systems. BERTScore correlates better with human judgments and provides stronger model selection performance than existing metrics. Finally, we use an adversarial paraphrase detection task to show that BE","claim_type":"abstract","evidence_strength":"source_metadata"},{"claim_text":"BERTScorecomputes semantic similarity between generated and reference ADRs based on contextualized embeddings from pretrained BERT models. Unlike surface-level metrics, BERTScore captures meaning and conceptual alignment even when wording differs. This makes it particularly suitable for evaluating architec- tural documentation where paraphrasing is common. [20]. BLEUmeasures n-gram precision to capture surface-level over- lap between generated and reference texts. Originally developed for machin","claim_type":"method","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"What remains inadequately explored is not merely the retrieval of more accurate information, but rather the ability to perform abstract synthesis based on retrieved evidence under controlled evaluation settings. This mismatch also poses an evaluation challenge for abstract QA. Conventional lexical and semantic metrics, such as ROUGE [17] and BERTScore [18], are too coarse-grained for long-form abstract answers, whose wording may vary substantially while differing in topical coverage. QA-based an","claim_type":"background","confidence":0.9,"evidence_strength":"citation_context"},{"claim_text":"The average rank improvement across all queries where adversarial modification is applied. This metric captures the magnitude of rank shifts introduced by adversarial edits. 5.4.2 Content Fidelity Metrics.Consistent with [ 56], to evaluate how well adversarial documents preserve the semantic and structural integrity of the original payload, we employ the following metrics: Semantic Similarity (SS):The average BERTScore F1 [ 58] between the original and adversarial documents, which measures token","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"For motion, we follow prior work [30,97] and report joint accuracy (MPJPE, PA-MPJPE), mesh accuracy (MPJVE), and trajectory error (MTE) for long-horizon prediction, all in millimeters, and joint rotation error (MPJRE) in degrees. For text, we adopt the linguistic metrics from MotionGPT3 [121] and report BLEU [65], ROUGE-L [53], CIDEr [81], and BERT-Score [114]. For 3D layout, we report 3D bounding box IoU (3D- IoU), precision and recall at IoU threshold 0.5 (P@0.5, R@0.5), and identity- level pr","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"We evaluate two aspects:spatial overlap, with Jaccard at K (J@K) and soft Recall (SR@K, the fraction of queries whose best Jaccard exceeds 0.8), and ranking quality, with hard Recall (R@K) and Mean Reciprocal Rank (MRR). Task 3: Trajectory Captioning.Given a trajectory, a model must generate a factual natural language caption. We evaluatelanguage qualitywith BERTScore F1 ( BS-F1) [31], ROUGE-L (R-L) [15], and METEOR [2], andspatial groundingis measured by POI Recall ( POI-R), the proportion of g","claim_type":"method","confidence":0.85,"evidence_strength":"citation_context"},{"claim_text":"Table 1 | Statistics summary of datasets by sampling bias. Detailed investigations of the ConcreteBatch are performed regarding sampling bias, hardness, and intra-modal similari- ties. CLIPScores are measured with a PE-Core-L-14-336 [46] pretrained on the MetaCLIP-5.4B dataset [47] (total 58B sam- ples seen), DINOv2 [48] is employed for DINOScores, and BERTScores [49] is calculated using the default RoBERTa- large [50] with baseline rescaling. For notational convenience, let Dℎ𝑐, D𝑙𝑐 and D𝑤𝑜 den","claim_type":"method","confidence":0.8,"evidence_strength":"citation_context"}],"why_cited":"Pith tracks BERTScore: Evaluating Text Generation with BERT because it crossed a citation-hub threshold. Current citing contexts most often use it as method evidence (7 contexts).","role_counts":[{"n":7,"context_role":"method"},{"n":6,"context_role":"background"},{"n":1,"context_role":"baseline"},{"n":1,"context_role":"other"}]},"error":null,"updated_at":"2026-05-20T06:41:48.716015+00:00"},"summary_claims":{"job_type":"summary_claims","status":"succeeded","result":{"title":"BERTScore: Evaluating Text Generation with BERT","claims":[{"claim_text":"We propose BERTScore, an automatic evaluation metric for text generation. Analogously to common metrics, BERTScore computes a similarity score for each token in the candidate sentence with each token in the reference sentence. However, instead of exact matches, we compute token similarity using contextual embeddings. We evaluate using the outputs of 363 machine translation and image captioning systems. BERTScore correlates better with human judgments and provides stronger model selection performance than existing metrics. Finally, we use an adversarial paraphrase detection task to show that BE","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks BERTScore: Evaluating Text Generation with BERT because it crossed a citation-hub threshold.","role_counts":[]},"error":null,"updated_at":"2026-05-14T09:08:01.858795+00:00"}},"summary":{"title":"BERTScore: Evaluating Text Generation with BERT","claims":[{"claim_text":"We propose BERTScore, an automatic evaluation metric for text generation. Analogously to common metrics, BERTScore computes a similarity score for each token in the candidate sentence with each token in the reference sentence. However, instead of exact matches, we compute token similarity using contextual embeddings. We evaluate using the outputs of 363 machine translation and image captioning systems. BERTScore correlates better with human judgments and provides stronger model selection performance than existing metrics. Finally, we use an adversarial paraphrase detection task to show that BE","claim_type":"abstract","evidence_strength":"source_metadata"}],"why_cited":"Pith tracks BERTScore: Evaluating Text Generation with BERT because it crossed a citation-hub threshold.","role_counts":[]},"graph":{"co_cited":[{"title":"The Llama 3 Herd of Models","work_id":"1549a635-88af-4ac1-acfe-51ae7bb53345","shared_citers":12},{"title":"Qwen3 Technical Report","work_id":"25a4e30c-1232-48e7-9925-02fa12ba7c9e","shared_citers":11},{"title":"Gemini 2.5: Pushing the Frontier with Advanced Reasoning, Multimodality, Long Context, and Next Generation Agentic Capabilities","work_id":"008df105-2fdd-45d8-857a-8e35868aecb6","shared_citers":10},{"title":"GPT-4 Technical Report","work_id":"b928e041-6991-4c08-8c81-0359e4097c7b","shared_citers":9},{"title":"Mistral 7B","work_id":"eb5e1305-ad11-4875-ad8d-ad8b8f697599","shared_citers":8},{"title":"BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding","work_id":"ed240a10-5b19-406c-baa5-30803f465785","shared_citers":7},{"title":"Qwen3-VL Technical Report","work_id":"1fe243aa-e3c0-4da6-b391-4cbcfc88d5c0","shared_citers":7},{"title":"DeepSeek-R1: Incentivizing Reasoning Capability in LLMs via Reinforcement Learning","work_id":"e6b75ad5-2877-4168-97c8-710407094d20","shared_citers":6},{"title":"Llama 2: Open Foundation and Fine-Tuned Chat Models","work_id":"68a5177f-d644-44c1-bd4f-4e5278c22f5d","shared_citers":6},{"title":"Qwen2.5-VL Technical Report","work_id":"69dffacb-bfe8-442d-be86-48624c60426f","shared_citers":6},{"title":"Bleurt: Learning robust metrics for text gener- ation","work_id":"55ab59c9-b06b-48c6-829e-64343cc28679","shared_citers":5},{"title":"Gemini: A Family of Highly Capable Multimodal Models","work_id":"83f7c85b-3f11-450f-ac0c-64d9745220b2","shared_citers":5},{"title":"LoRA: Low-Rank Adaptation of Large Language Models","work_id":"0426219a-789e-4964-adc8-a04538510818","shared_citers":5},{"title":"RoBERTa: A Robustly Optimized BERT Pretraining Approach","work_id":"41fe12c4-e538-4890-a244-480650ed3078","shared_citers":5},{"title":"Sentence-BERT: Sentence Embeddings using Siamese BERT-Networks","work_id":"27adfcc9-2a67-43d6-a844-78309012411f","shared_citers":5},{"title":"DeepSeekMath: Pushing the Limits of Mathematical Reasoning in Open Language Models","work_id":"c5006563-f3ec-438a-9e35-b7b484f34828","shared_citers":4},{"title":"Gemma 3 Technical Report","work_id":"f93e08bf-9e96-409b-8ac6-b8385fd17fd7","shared_citers":4},{"title":"Gemma: Open Models Based on Gemini Research and Technology","work_id":"a9ea2870-df28-40b8-a9e0-a7e9a116f793","shared_citers":4},{"title":"G-Eval: NLG Evaluation using GPT-4 with Better Human Alignment","work_id":"6bc7f654-ec90-4d95-b023-416717a28f45","shared_citers":4},{"title":"Gptscore: Evaluate as you desire.arXiv preprint arXiv:2302.04166","work_id":"13d9a137-d114-4c30-aa62-3b8d25d9eb7c","shared_citers":4},{"title":"InternVL3.5: Advancing Open-Source Multimodal Models in Versatility, Reasoning, and Efficiency","work_id":"b8f5e260-fff5-444e-bcf5-2c42cfefd83d","shared_citers":4},{"title":"Judging LLM-as-a-Judge with MT-Bench and Chatbot Arena","work_id":"d0c30cd7-81e1-4159-a87f-f6adca77ff08","shared_citers":4},{"title":"LLaMA: Open and Efficient Foundation Language Models","work_id":"c018fc23-6f3f-4035-9d02-28a2173b2b9d","shared_citers":4},{"title":"Qwen3 Embedding: Advancing Text Embedding and Reranking Through Foundation Models","work_id":"bab684a8-d933-426c-a19e-2c855a0d1f59","shared_citers":4}],"time_series":[{"n":1,"year":2019},{"n":1,"year":2023},{"n":1,"year":2024},{"n":66,"year":2026}],"dependency_candidates":[]},"authors":[{"id":"31d5a154-a832-4674-afc0-4a3d03df02f4","orcid":null,"display_name":"Felix Wu","source":"manual","import_confidence":0.72},{"id":"a6bd0858-91d6-47df-8ce5-e08b05941e15","orcid":null,"display_name":"Kilian Q. Weinberger","source":"manual","import_confidence":0.72},{"id":"42670f71-d0ea-4278-9b4e-2046cc00cf9e","orcid":null,"display_name":"Tianyi Zhang","source":"manual","import_confidence":0.72},{"id":"b627ff25-180a-48fd-bc53-854cd6654549","orcid":null,"display_name":"Varsha Kishore","source":"manual","import_confidence":0.72},{"id":"1a505e27-523b-4a50-bb24-88475db88b78","orcid":null,"display_name":"Yoav Artzi","source":"manual","import_confidence":0.72}]}}