Tempus delivers 607 GOPS at 10.677 W using fixed 16 AIE cores on Versal AI Edge, with 211.2x better platform-aware utility than spatial SOTA ARIES and zero URAM/DSP utilization.
BERT: a review of applications in natural language processing and understanding
12 Pith papers cite this work. Polarity classification is still indexing.
representative citing papers
SAri-RFT applies GRPO-based reinforcement fine-tuning to LVLMs on novel two-term and three-term visual semantic arithmetic tasks, reaching SOTA on the new IRPD dataset and Visual7W-Telling.
SELF-EMO lets LLMs bootstrap better emotion recognition and expression via self-play, data flywheel filtering with smoothed IoU rewards, and SELF-GRPO reinforcement learning, yielding SOTA gains on IEMOCAP, MELD, and EmoryNLP.
A physics-informed Fourier-wavelet transformer model reports the lowest normalized mean-squared error on cylinder-wake and fluid-structure interaction velocity-field benchmarks compared with spectral, transformer, operator-learning, and PINN baselines.
AFIP is a training-free attention-correction method that cuts object hallucination rates in multimodal LLMs by concentrating cross-head attention and restoring faded visual attention.
A transformer framework for composed vision-language retrieval in skin cancer uses hierarchical query representations and global-local alignment to improve performance over prior methods on the Derm7pt dataset.
A new dataset of 400k visual instructions including negative examples at three semantic levels reduces hallucinations in models like MiniGPT-4 when used for fine-tuning while improving benchmark performance.
RTP-LLM is a new LLM inference engine achieving 4.7x-6.3x model loading speedup and 1.12x-2.52x throughput gains over vLLM and SGLang via disaggregated phases, multi-tier KV cache, and modular optimizations in production at Alibaba.
CS-PQ optimizes PQ construction on CPUs via vectorized SIMD across centroids and cache-restructured pipelines, claiming 10.7x speedup over prior CPU methods with no accuracy loss on large datasets.
Hardware approximations for Softmax and LayerNorm preserve exact normalization guarantees and deliver up to 14x area reduction in 28nm silicon with negligible accuracy loss on GLUE, SQuAD, and perplexity.
THInfer achieves 62-84% higher throughput than GPU baselines for Llama 7B-30B models on MT-3000 through bandwidth-focused co-design, and runs 70B models where GPU frameworks fail.
LLMs perform well on basic syntactic and semantic bugs in small code but struggle with complex security vulnerabilities and large production codebases.
citing papers explorer
-
Tempus: A Temporally Scalable Resource-Invariant GEMM Streaming Framework for Versal AI Edge
Tempus delivers 607 GOPS at 10.677 W using fixed 16 AIE cores on Versal AI Edge, with 211.2x better platform-aware utility than spatial SOTA ARIES and zero URAM/DSP utilization.
-
Multi-modal Reasoning with LLMs for Visual Semantic Arithmetic
SAri-RFT applies GRPO-based reinforcement fine-tuning to LVLMs on novel two-term and three-term visual semantic arithmetic tasks, reaching SOTA on the new IRPD dataset and Visual7W-Telling.
-
SELF-EMO: Emotional Self-Evolution from Recognition to Consistent Expression
SELF-EMO lets LLMs bootstrap better emotion recognition and expression via self-play, data flywheel filtering with smoothed IoU rewards, and SELF-GRPO reinforcement learning, yielding SOTA gains on IEMOCAP, MELD, and EmoryNLP.
-
A Physics-Informed Fourier-Wavelet Transformer for Multiscale Computational Fluid Dynamics Surrogate Modeling
A physics-informed Fourier-wavelet transformer model reports the lowest normalized mean-squared error on cylinder-wake and fluid-structure interaction velocity-field benchmarks compared with spectral, transformer, operator-learning, and PINN baselines.
-
Correcting Visual Blur Induced by Attention Distraction to Reduce Hallucinations: Algorithm and Theory
AFIP is a training-free attention-correction method that cuts object hallucination rates in multimodal LLMs by concentrating cross-head attention and restoring faded visual attention.
-
Composed Vision-Language Retrieval for Skin Cancer Case Search via Joint Alignment of Global and Local Representations
A transformer framework for composed vision-language retrieval in skin cancer uses hierarchical query representations and global-local alignment to improve performance over prior methods on the Derm7pt dataset.
-
Mitigating Hallucination in Large Multi-Modal Models via Robust Instruction Tuning
A new dataset of 400k visual instructions including negative examples at three semantic levels reduces hallucinations in models like MiniGPT-4 when used for fine-tuning while improving benchmark performance.
-
RTP-LLM: High-Performance Alibaba LLM Inference Engine
RTP-LLM is a new LLM inference engine achieving 4.7x-6.3x model loading speedup and 1.12x-2.52x throughput gains over vLLM and SGLang via disaggregated phases, multi-tier KV cache, and modular optimizations in production at Alibaba.
-
CS-PQ: Cache-Friendly SIMD Product Quantization for Large-Scale ANNS Index Construction
CS-PQ optimizes PQ construction on CPUs via vectorized SIMD across centroids and cache-restructured pipelines, claiming 10.7x speedup over prior CPU methods with no accuracy loss on large datasets.
-
Hardware-Efficient Softmax and Layer Normalization with Guaranteed Normalization for Edge Devices
Hardware approximations for Softmax and LayerNorm preserve exact normalization guarantees and deliver up to 14x area reduction in 28nm silicon with negligible accuracy loss on GLUE, SQuAD, and perplexity.
-
Bandwidth-Aware LLM Inference on Heterogeneous Many-Core Supercomputers
THInfer achieves 62-84% higher throughput than GPU baselines for Llama 7B-30B models on MT-3000 through bandwidth-focused co-design, and runs 70B models where GPU frameworks fail.
-
Can LLMs Find Bugs in Code? An Evaluation from Beginner Errors to Security Vulnerabilities in Python and C++
LLMs perform well on basic syntactic and semantic bugs in small code but struggle with complex security vulnerabilities and large production codebases.