REVIEW 3 cited by
MaLei at the PLABA Track of TREC 2024: RoBERTa for Term Replacement -- LLaMA3.1 and GPT-4o for Complete Abstract Adaptation
Not yet reviewed by Pith; the record is open.
This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.
SPECIMEN: schema-true, not a live event
T0 review · schema-true
One-sentence machine reading of the paper's core claim.
pith:XXXXXXXX · record.json · timestamp
read the original abstract
This report is the system description of the MaLei team (Manchester and Leiden) for the shared task Plain Language Adaptation of Biomedical Abstracts (PLABA) 2024 (we had an earlier name BeeManc following last year), affiliated with TREC2024 (33rd Text REtrieval Conference https://ir.nist.gov/evalbase/conf/trec-2024). This report contains two sections corresponding to the two sub-tasks in PLABA-2024. In task one (term replacement), we applied fine-tuned ReBERTa-Base models to identify and classify the difficult terms, jargon, and acronyms in the biomedical abstracts and reported the F1 score (Task 1A and 1B). In task two (complete abstract adaptation), we leveraged Llamma3.1-70B-Instruct and GPT-4o with the one-shot prompts to complete the abstract adaptation and reported the scores in BLEU, SARI, BERTScore, LENS, and SALSA. From the official Evaluation from PLABA-2024 on Task 1A and 1B, our much smaller fine-tuned RoBERTa-Base model ranked 3rd and 2nd respectively on the two sub-tasks, and the 1st on averaged F1 scores across the two tasks from 9 evaluated systems. Our LLaMA-3.1-70B-instructed model achieved the highest Completeness score for Task 2. We share our source codes, fine-tuned models, and related resources at https://github.com/HECTA-UoM/PLABA2024
Forward citations
Cited by 3 Pith papers
-
Lessons from the TREC Plain Language Adaptation of Biomedical Abstracts (PLABA) track
Across two TREC shared-task years, top LLM systems matched human writers on factual accuracy and completeness but not on simplicity or brevity, while common automatic metrics correlated poorly with manual judgments.
-
MaLei at MultiClinSUM: Summarisation of Clinical Documents using Perspective-Aware Iterative Self-Prompting with LLMs
Perspective-aware iterative self-prompting with GPT-4o summarizes clinical case reports with high semantic fidelity (BERTScore F1 0.85) but low lexical overlap (ROUGE-L F1 0.31).
-
DLP: Dynamic Layerwise Pruning in Large Language Models
DLP assigns each LLM layer a sparsity rate derived from the median of Wanda-style weight-activation scores, improving perplexity and zero-shot accuracy at high sparsity versus uniform and outlier-based layerwise pruning.
Discussion (0). Sign in to comment.