Pith. sign in

Paper Citation Record · LEDGER

FaceInsight: A Multimodal Large Language Model for Face Perception

As of 18 August 2026, this Paper Citation Record lists 58 of 58 outbound references and 0 inbound Pith citation observations for arXiv:2504.15624.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2504.15624 v2

Coverage vector

measured 58 of 58 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-16T11:24:53.697821Z

measured 58 of 58 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-18T06:34:40.430872+00:00

measured 0 of 0 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

58 of 58 outbound references displayed

  • verified exact2
  • verified fuzzy26
  • unresolved30
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 003f1676-ad98-47bf-b6c9-d51a5e1ec0f4 · outbound

This paper cites Facexbench: Evaluating multimodal llms on face understanding.

FaceInsight: A Multimodal Large Language Model for Face Perception Facexbench: Evaluating multimodal llms on face understanding

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.508928Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.508928Z digest=sha256:3f08f5f30c0f36c57c4bae5a284699806bcc0437cc50d03d575c5d059bc91206

Observation c55908f8-8390-4a32-8b9c-956a993c15b6 · outbound

This paper cites Face-MLLM: A Large Face Perception Model.

FaceInsight: A Multimodal Large Language Model for Face Perception Face-MLLM: A Large Face Perception Model

Reference 2

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.513242Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.513242Z digest=sha256:88b1093d9f7cb8ab95765209385d8916cd994bb7a209a2a6b07ece37484a809e

Observation cbd471f6-cbb9-4ff3-a666-acbf55b4a1f4 · outbound

This paper cites FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO.

FaceInsight: A Multimodal Large Language Model for Face Perception FaVChat: Hierarchical Prompt-Query Guided Facial Video Understanding with Data-Efficient GRPO

Reference 3

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.517329Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.517329Z digest=sha256:4d7d351accddc37a19e8c63bb2d99261768a025d19f3b28d6895986208c4672e

Observation 70c70e5f-92d2-40f5-a92e-1b3dadbbd9d4 · outbound

This paper cites EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning.

FaceInsight: A Multimodal Large Language Model for Face Perception EMO-LLaMA: Enhancing Facial Emotion Understanding with Instruction Tuning

Reference 4

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.521271Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.521271Z digest=sha256:bee4f3b98f5ccba5ed886a415589a715231287a0e89bf36cfced9aa6601c8587

Observation 16743285-a98a-481c-8fcd-47cf3bc3ae4a · outbound

This paper cites OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models.

FaceInsight: A Multimodal Large Language Model for Face Perception OpenFlamingo: An Open-Source Framework for Training Large Autoregressive Vision-Language Models

Reference 5

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.525193Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.525193Z digest=sha256:79902ebede2626f50d2a14dc4450912d0760efc9b82aafa051ba4c8e28cdc711

Observation ac7b534f-4c34-4db2-8616-3313911510ae · outbound

This paper cites Visual instruction tuning.

FaceInsight: A Multimodal Large Language Model for Face Perception Visual instruction tuning

Reference 6

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.305708Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.529154Z digest=sha256:bafc24e3b1533ab98d11114f48f26daa52984420e3a688bde018a5f42ec6ad29

Observation 4075910d-50fb-47de-ac7e-a18264d5e592 · outbound

This paper cites Llava-next: Improved reasoning, ocr, and world knowledge, 2024.

FaceInsight: A Multimodal Large Language Model for Face Perception Llava-next: Improved reasoning, ocr, and world knowledge, 2024

Reference 7

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.532784Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.532784Z digest=sha256:ebd31bf9242909571cb5ff587e9b6d62bed4582e5b696f3241b1add6aad805aa

Observation 46013e8b-f1a1-4a6f-93f7-2a2aad8421d5 · outbound

This paper cites Visual Large Language Models for Generalized and Specialized Applications.

FaceInsight: A Multimodal Large Language Model for Face Perception Visual Large Language Models for Generalized and Specialized Applications

Reference 8

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.536150Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.536150Z digest=sha256:4b40eb4cbe7a47aae6a91549d3013daead7e4a375114a580f6ea64728fed64d5

Observation 2edbff60-3d85-4a88-b8fb-da8af66dae09 · outbound

This paper cites Learning transferable visual models from natural language supervision.

FaceInsight: A Multimodal Large Language Model for Face Perception Learning transferable visual models from natural language supervision

Reference 9

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.539671Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.539671Z digest=sha256:f0c5970d42c0a171d68a54b8ced20e7bdf1603dffa0cc9c79a83e4d4268a570f

Observation 87d35aa3-53c4-4611-8b07-f5f83b08715e · outbound

This paper cites Vcoder: Versatile vision encoders for multimodal large language models.

FaceInsight: A Multimodal Large Language Model for Face Perception Vcoder: Versatile vision encoders for multimodal large language models

Reference 10

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.284888Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.542884Z digest=sha256:89648e138eda02c47e7e34f9bb2c47c5877a1f1c5edb90d63886401fb551b490

Observation c34e4181-6768-4dd6-9df0-0bd0d559b1cb · outbound

This paper cites Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality.

FaceInsight: A Multimodal Large Language Model for Face Perception Vicuna: An open-source chatbot impressing gpt-4 with 90%* chatgpt quality

Reference 11

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.546187Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.546187Z digest=sha256:42fa0d4c7ab74383323f5e557ba0392283c4deb32ffbfdc8495ea6bbd8f5e125

Observation 2240f453-8216-4db9-a520-a9bec520dec5 · outbound

This paper cites CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning.

FaceInsight: A Multimodal Large Language Model for Face Perception CogCoM: A Visual Language Model with Chain-of-Manipulations Reasoning

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.549430Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.549430Z digest=sha256:fde0c55c1056b78255597656c5f9e611e64d5e3ad5f90523505f0ae2c3a9f648

Observation 544ef856-1577-49c6-b3fb-5285396eb874 · outbound

This paper cites LLaMA: Open and Efficient Foundation Language Models.

FaceInsight: A Multimodal Large Language Model for Face Perception LLaMA: Open and Efficient Foundation Language Models

Reference 13

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.552822Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.552822Z digest=sha256:0e87a32aa98fc0ebc64c4fc71f6ba7371a57f90ca82054688971ddef9193943e

Observation cbc2656a-0d78-421c-91a9-48172c79843a · outbound

This paper cites MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models.

FaceInsight: A Multimodal Large Language Model for Face Perception MiniGPT-4: Enhancing Vision-Language Understanding with Advanced Large Language Models

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.556124Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.556124Z digest=sha256:af75f625330f6af072399dfa79fcbd254bdc80bce42bd324f032d0f5b7416d5b

Observation 585e1b60-d5ec-4824-8e8d-33c0c60eea72 · outbound

This paper cites Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models.

FaceInsight: A Multimodal Large Language Model for Face Perception Blip-2: Bootstrapping language-image pre-training with frozen image encoders and large language models

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.559291Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.559291Z digest=sha256:8b6403b5bc66e8726a03ccacabdcd3406fb88423130dd14fe22e5582228102d2

Observation f2330d3b-d5ff-4178-8b4e-383aaf5573be · outbound

This paper cites Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023.

FaceInsight: A Multimodal Large Language Model for Face Perception Instructblip: Towards general-purpose vision-language models with instruction tuning, 2023

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.562392Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.562392Z digest=sha256:1ef9b7d8e5ef777cd20005024150f46a3eb9c40bfb5c862f4dd19fbf575b4366

Observation 7a28999c-b430-4793-a13b-0cacc6200e2a · outbound

This paper cites mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality.

FaceInsight: A Multimodal Large Language Model for Face Perception mPLUG-Owl: Modularization Empowers Large Language Models with Multimodality

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.565583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.565583Z digest=sha256:4731e1cbc478c6747bc11c21d59dfc059cf16f9abb6813d37fdf74cbaa46350e

Observation ef6e2b7f-2c86-4f95-ac15-4edd82480d45 · outbound

This paper cites Improved baselines with visual instruction tuning.

FaceInsight: A Multimodal Large Language Model for Face Perception Improved baselines with visual instruction tuning

Reference 18

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.569251Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.569251Z digest=sha256:5d122096c6c0909e83c32d6c7e3055b4d19bc8e482bc91a03fe198586b146930

Observation d43aed7b-a458-46b0-b326-1e90910b3f36 · outbound

This paper cites Lion: Empowering multimodal large language model with dual-level visual knowl- edge.

FaceInsight: A Multimodal Large Language Model for Face Perception Lion: Empowering multimodal large language model with dual-level visual knowl- edge

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.252144Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.572636Z digest=sha256:945dde80dd408bace1e4b381ef1273d86cbef75722d6ed6ffc8b35d25a7cc4bc

Observation 092a2334-530a-44d1-a585-d7f5a859f5c5 · outbound

This paper cites MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning.

FaceInsight: A Multimodal Large Language Model for Face Perception MiniGPT-v2: large language model as a unified interface for vision-language multi-task learning

Reference 20

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.575711Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.575711Z digest=sha256:782fa60180b1e385459a5e94c8b47a5b8e845bc29bcf8dd502aec93e0569c0f6

Observation dfc45496-c8f6-426d-831b-6e383433b020 · outbound

This paper cites Vila: On pre-training for visual language models.

FaceInsight: A Multimodal Large Language Model for Face Perception Vila: On pre-training for visual language models

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.242612Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.579312Z digest=sha256:ce82c4fd59ce8d3cafb73af73e806dec2b577d46ae9bb1a9228cdad08e192daa

Observation 5f74f07f-2cab-4692-9b80-f10acccb1f34 · outbound

This paper cites MoE-LLaVA: Mixture of Experts for Large Vision-Language Models.

FaceInsight: A Multimodal Large Language Model for Face Perception MoE-LLaVA: Mixture of Experts for Large Vision-Language Models

Reference 22

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.582221Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.582221Z digest=sha256:78685892e2235eb78569ef0fa10a93fe7b8a223c1cb4eb0e3f13c99684affd8f

Observation aa924b58-fb9d-4292-a756-9f4c8c8c4f2f · outbound

This paper cites Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling.

FaceInsight: A Multimodal Large Language Model for Face Perception Expanding Performance Boundaries of Open-Source Multimodal Models with Model, Data, and Test-Time Scaling

Reference 23

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.585558Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.585558Z digest=sha256:2e1b5db0330111727c54070af98d812e5b6a29efa3742fc4c916a79d0006d7bc

Observation 83f625b7-d49c-4d49-9556-47597e606d4d · outbound

This paper cites MiniCPM-V: A GPT-4V Level MLLM on Your Phone.

FaceInsight: A Multimodal Large Language Model for Face Perception MiniCPM-V: A GPT-4V Level MLLM on Your Phone

Reference 24

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.588905Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.588905Z digest=sha256:93901d097fcdff3b6fec9ea4c91f0e5159f459b16d724f365388c80fb37b0bf9

Observation 47f935ce-e833-4e43-b5d3-7b3c328e7133 · outbound

This paper cites MANTIS: Interleaved Multi-Image Instruction Tuning.

FaceInsight: A Multimodal Large Language Model for Face Perception MANTIS: Interleaved Multi-Image Instruction Tuning

Reference 25

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.592303Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.592303Z digest=sha256:6357114d267e46a143cb80491dd70b9aa088738ad4b0f048c2301bc4e1b2ebdc

Observation 2fc529ec-1a6e-4d32-87fc-fbe7e1a2f4ce · outbound

This paper cites Monkey: Image resolution and text label are important things for large multi-modal models.

FaceInsight: A Multimodal Large Language Model for Face Perception Monkey: Image resolution and text label are important things for large multi-modal models

Reference 26

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.233439Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.595487Z digest=sha256:9647fd955af4acb43ed271c96d222e1f610997c86bda4649f67496f13e3c1c03

Observation c70f398e-67df-45db-9aff-f9a824ac63d9 · outbound

This paper cites Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone.

FaceInsight: A Multimodal Large Language Model for Face Perception Phi-3 Technical Report: A Highly Capable Language Model Locally on Your Phone

Reference 27

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.598754Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.598754Z digest=sha256:205a0d553ad8d591e2b6a48bc21460a6c86e9d512857f72a26534876a324fb15

Observation 688818ff-96ec-4e24-ba0e-18d65a22de72 · outbound

This paper cites Qwen2.5-VL Technical Report.

FaceInsight: A Multimodal Large Language Model for Face Perception Qwen2.5-VL Technical Report

Reference 28

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.602074Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.602074Z digest=sha256:0b8c918555d3d225d225c9b6e891f67b382b7fbb9e2d3810afe6463828097971

Observation dea68062-72c4-405f-8db9-8744b187c9e7 · outbound

This paper cites Deepveil: deep learning for identification of face, gender, expression recognition under veiled conditions.

FaceInsight: A Multimodal Large Language Model for Face Perception Deepveil: deep learning for identification of face, gender, expression recognition under veiled conditions

Reference 29

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.223774Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.605335Z digest=sha256:6f8de90f94061d34d2e758b9aacee7483916b26f6ef6f108cf409a36a939313e

Observation 294468b2-398c-46e8-9cb2-117051d19e51 · outbound

This paper cites Spl- net: Spatial-semantic patch learning network for facial attribute recognition with limited labeled data.

FaceInsight: A Multimodal Large Language Model for Face Perception Spl- net: Spatial-semantic patch learning network for facial attribute recognition with limited labeled data

Reference 30

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.214140Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.608240Z digest=sha256:a9afb9117740c27485370756cd62600c88954ae1068ed58a6151ae7b187f08ad

Observation 2d2a09aa-0ff2-4b83-be11-fbb0cea23e00 · outbound

This paper cites Logical consistency and greater descriptive power for facial hair attribute learning.

FaceInsight: A Multimodal Large Language Model for Face Perception Logical consistency and greater descriptive power for facial hair attribute learning

Reference 31

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.204474Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.611260Z digest=sha256:a1c46cfe78b2b7bf9083d7a57618ef1e4e82614e19856d985451e9604bf835c0

Observation a3cd4719-bf61-4aa8-adaf-9bb01a8550ba · outbound

This paper cites LogicNet: A Logical Consistency Embedded Face Attribute Learning Network.

FaceInsight: A Multimodal Large Language Model for Face Perception LogicNet: A Logical Consistency Embedded Face Attribute Learning Network

Reference 32

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:24:53.781617Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.614264Z digest=sha256:047a71e11c8b272b6d5dc54383b7fd3b513b7de0e2d489f73e758f4939a8cd85

Observation 46f84d71-a489-4c1c-b696-7e3222d3e90a · outbound

This paper cites A survey on facial emotion recognition techniques: A state-of-the-art literature review.

FaceInsight: A Multimodal Large Language Model for Face Perception A survey on facial emotion recognition techniques: A state-of-the-art literature review

Reference 33

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.193361Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.617519Z digest=sha256:6aa77a7a94fd357f3b33a27a7e7ba0c2c2b78320124fcd8198494e43b02f05af

Observation f8694d9f-42f0-4c95-a6d7-39fd5384e972 · outbound

This paper cites Semi-supervised multimodal emotion recognition with expression mae.

FaceInsight: A Multimodal Large Language Model for Face Perception Semi-supervised multimodal emotion recognition with expression mae

Reference 34

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.183340Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.620720Z digest=sha256:38b959ba4db5feb249b7969d466fcd3f065f9f94e5c6b4b50e69ae92dcc046ee

Observation 85958bd2-4e15-40b6-a613-541e04e46479 · outbound

This paper cites Facial affective behavior analysis with instruction tuning.

FaceInsight: A Multimodal Large Language Model for Face Perception Facial affective behavior analysis with instruction tuning

Reference 35

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.173959Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.623918Z digest=sha256:443707f02efded9595a5681849e4cf91cc30ba73e2bfd8f91df445b305f80c24

Observation 600035e0-bcc6-4c88-9bf3-c4a140d7dd51 · outbound

This paper cites Rank consistent ordinal regression for neural networks with application to age estimation.

FaceInsight: A Multimodal Large Language Model for Face Perception Rank consistent ordinal regression for neural networks with application to age estimation

Reference 36

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.164520Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.627056Z digest=sha256:41b591de0973d1c3b0befc6e5881bdc94d8829f6d17e3897573ebddb4c77f7d1

Observation 5b11fdbf-9eba-4ba5-bd24-a68dfb55f365 · outbound

This paper cites Learning probabilistic ordinal embeddings for uncertainty-aware regression.

FaceInsight: A Multimodal Large Language Model for Face Perception Learning probabilistic ordinal embeddings for uncertainty-aware regression

Reference 37

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.154732Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.630070Z digest=sha256:cbd286a76b057b05987e5701545624e048bb55634e0e183c4f865d94aeea42b1

Observation 7c258d5d-b3d1-46c9-aa58-a3c01c25e9e9 · outbound

This paper cites Mivolo: Multi-input transformer for age and gender estimation.

FaceInsight: A Multimodal Large Language Model for Face Perception Mivolo: Multi-input transformer for age and gender estimation

Reference 38

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.144228Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.633220Z digest=sha256:5dfa932e7708a76892db5b6bed8f28f2392559666828dfacbbf7a6abcd952822

Observation 09de563e-7411-4cf6-8c47-55cad912aaa1 · outbound

This paper cites Hyperface: A deep multi-task learning framework for face detection, landmark localization, pose estimation, and gender recognition.

FaceInsight: A Multimodal Large Language Model for Face Perception Hyperface: A deep multi-task learning framework for face detection, landmark localization, pose estimation, and gender recognition

Reference 39

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.636220Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.636220Z digest=sha256:1fc467ec34fcb2dc9c82ffff4d52b7f95881d46366bb49410c8b59265862d1eb

Observation bb8fb21e-56f1-4813-afad-4b20e01d60f4 · outbound

This paper cites An all-in-one convolutional neural network for face analysis.

FaceInsight: A Multimodal Large Language Model for Face Perception An all-in-one convolutional neural network for face analysis

Reference 40

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.127688Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.639265Z digest=sha256:011d219cbe07f0d79cde82d0f5812d0978b39e8caee9bb191dfc5e4504528e55

Observation 52f0af8c-fcb9-4e13-b428-b4f283db4b71 · outbound

This paper cites Swinface: a multi-task transformer for face recognition, expression recog- nition, age estimation and attribute estimation.

FaceInsight: A Multimodal Large Language Model for Face Perception Swinface: a multi-task transformer for face recognition, expression recog- nition, age estimation and attribute estimation

Reference 41

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.117458Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.642253Z digest=sha256:66d56526fa705dcf3b68a15b26218fc12524da5950056ab6b557238b5180cf59

Observation fb55f8b9-47e2-4037-a952-49f808459d89 · outbound

This paper cites FaceXFormer: A Unified Transformer for Facial Analysis.

FaceInsight: A Multimodal Large Language Model for Face Perception FaceXFormer: A Unified Transformer for Facial Analysis

Reference 42

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.645333Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.645333Z digest=sha256:ab351b08cb548a61fa2835b89cdbf542f7d9048d69537e84eb95577158258ef0

Observation 43e2e7b7-187a-43aa-9f2a-b4f65d6778c3 · outbound

This paper cites Task-adaptive Q-Face.

FaceInsight: A Multimodal Large Language Model for Face Perception Task-adaptive Q-Face

Reference 43

Resolution
verified exact
local_arxiv, observed 2026-08-16T11:24:53.758113Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.648787Z digest=sha256:eac98c472c1edbbf2325d5f74d5181fc39c3cb88de15014b3fa98f3996e42540

Observation d437272b-b79c-4012-9640-ac647777b30e · outbound

This paper cites Faceptor: A generalist model for face perception.

FaceInsight: A Multimodal Large Language Model for Face Perception Faceptor: A generalist model for face perception

Reference 44

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.106799Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.652112Z digest=sha256:5f5fee81827f672a834b83aa942652c09043bbf6481b041fccefcd473a818cbd

Observation bd779c3d-f731-4559-90ad-299ee1948a9d · outbound

This paper cites General facial repre- sentation learning in a visual-linguistic manner.

FaceInsight: A Multimodal Large Language Model for Face Perception General facial repre- sentation learning in a visual-linguistic manner

Reference 45

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.096728Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.655321Z digest=sha256:f2c27bce5fdcdc73f190574676357288df57b1dc905296dbb6bb1e91cdad8fe7

Observation 4b4b9eb3-9661-4306-be4e-acd5399ff10e · outbound

This paper cites Pre-training strategies and datasets for facial representation learning.

FaceInsight: A Multimodal Large Language Model for Face Perception Pre-training strategies and datasets for facial representation learning

Reference 46

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.086240Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.658443Z digest=sha256:ae293c5d26a2cdc3d1d4fd66ccfbbe2132aad02ae8b7014170960b91bb77aa80

Observation 88648c52-8176-4120-ae86-e34ae78e2635 · outbound

This paper cites Label2label: A language modeling framework for multi-attribute learning.

FaceInsight: A Multimodal Large Language Model for Face Perception Label2label: A language modeling framework for multi-attribute learning

Reference 47

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.076189Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.661529Z digest=sha256:19c5e87c5deccec77df09e400a67d7159e06bf872ec67d05d1c00b82065daa5e

Observation 9515983b-4c42-4782-8dce-afc340cde691 · outbound

This paper cites Prompting Visual-Language Models for Dynamic Facial Expression Recognition.

FaceInsight: A Multimodal Large Language Model for Face Perception Prompting Visual-Language Models for Dynamic Facial Expression Recognition

Reference 48

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.664735Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.664735Z digest=sha256:746607d88574bcf2c705aa61ac4752d28026d3f06b2ddf9fbcdcd44b295c3b06

Observation 4d431e98-8a95-426b-9ded-e41bdc4803c4 · outbound

This paper cites Emoclip: A vision-language method for zero-shot video facial expression recognition.

FaceInsight: A Multimodal Large Language Model for Face Perception Emoclip: A vision-language method for zero-shot video facial expression recognition

Reference 49

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.066359Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.668204Z digest=sha256:527d50de9c2a37862b4e692864b56701b3a5296ac1f5a67ec1dea95ad2576495

Observation 8da207b4-a2c3-42b2-8b5f-1a1e30d27f98 · outbound

This paper cites Retinaface: Single-shot multi-level face localisation in the wild.

FaceInsight: A Multimodal Large Language Model for Face Perception Retinaface: Single-shot multi-level face localisation in the wild

Reference 50

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.056787Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.671306Z digest=sha256:072b1968cd55005de69faecae2e084bde06bfa263a2a796dd470940f07e11127

Observation 2fbdaba1-7739-46d5-8c50-a6eb290f0146 · outbound

This paper cites Maad-face: A massively annotated attribute dataset for face images.

FaceInsight: A Multimodal Large Language Model for Face Perception Maad-face: A massively annotated attribute dataset for face images

Reference 51

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.047060Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.674939Z digest=sha256:10e13254db625a2e0e40dbfb92497b8c20b939bc89cba53a41a32bcdc58da88e

Observation 8b0f727b-f72f-45b4-95ba-f3e54be6c04c · outbound

This paper cites Deep learning face attributes in the wild.

FaceInsight: A Multimodal Large Language Model for Face Perception Deep learning face attributes in the wild

Reference 52

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.678340Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.678340Z digest=sha256:d61b3d84cef4d78de07cc2128a4eed0f80a169d5001044c1da01cdee0fc51897

Observation e78ecc66-61c9-4768-aad5-573b583d5b0d · outbound

This paper cites Fairface: Face attribute dataset for bal- anced race, gender, and age for bias measurement and mitigation.

FaceInsight: A Multimodal Large Language Model for Face Perception Fairface: Face attribute dataset for bal- anced race, gender, and age for bias measurement and mitigation

Reference 53

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.031247Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.681503Z digest=sha256:c35c5d58e50c4d9d9d8e8997d85f5c44acd3093c5c6de401bf755fff4ef2213d

Observation 8327ca4f-d382-44c7-a5c3-d3645559ba8c · outbound

This paper cites Age progression/regression by con- ditional adversarial autoencoder.

FaceInsight: A Multimodal Large Language Model for Face Perception Age progression/regression by con- ditional adversarial autoencoder

Reference 54

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.021080Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.684816Z digest=sha256:e45c504d5fdd5950451d2caedae4a787565d6c4b754f365562c837de48087253

Observation 9c0a7961-8e69-4445-8c39-1bd293f82873 · outbound

This paper cites From facial expression recognition to interpersonal relation prediction.

FaceInsight: A Multimodal Large Language Model for Face Perception From facial expression recognition to interpersonal relation prediction

Reference 55

Resolution
verified fuzzy
raw_fallback, observed 2026-08-16T11:24:54.011284Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-18T06:34:40.430872+00:00.

source=pdf_text observed=2026-08-16T11:24:53.688135Z digest=sha256:e5f3cc818304e79bcd33cf771006a808b75ddbdb603baf7cb3719bbb6b1f914a

Observation 6a6cf00c-837f-4d91-860f-405e9671cc11 · outbound

This paper cites Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild.

FaceInsight: A Multimodal Large Language Model for Face Perception Reliable crowdsourcing and deep locality-preserving learning for expression recognition in the wild

Reference 56

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.691170Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.691170Z digest=sha256:ddfba8b56e60627850e87893d9d64cd3b6b1600fdedcb4fa50ad3aea114848cb

Observation dba4ae96-a1e2-4302-86c5-a641927a8ac5 · outbound

This paper cites LLaVA-OneVision: Easy Visual Task Transfer.

FaceInsight: A Multimodal Large Language Model for Face Perception LLaVA-OneVision: Easy Visual Task Transfer

Reference 57

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.694415Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.694415Z digest=sha256:7a1d7373a2d45ac951324ba8d8a1db5a281f2726a7ca1bb56e04c6703b53748a

Observation be532ca1-a559-4c2b-a969-e491e677ddb0 · outbound

This paper cites NVILA: Efficient Frontier Visual Language Models.

FaceInsight: A Multimodal Large Language Model for Face Perception NVILA: Efficient Frontier Visual Language Models

Reference 58

Resolution
unresolved
no resolver link, observed 2026-08-16T11:24:53.697821Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-16T11:24:53.697821Z digest=sha256:4b8a831d20b122eb0fb608464d9872c7e8069f98398f396dd81f09cb6d74bf8b

Pith citing papers

No inbound Pith citation observations are available.