{"record_type":"pith_number_record","schema_url":"https://pith.science/schemas/pith-number/v1.json","pith_number":"pith:2024:6LF7IY6F3TF6IEVP76OH4ENLI4","short_pith_number":"pith:6LF7IY6F","schema_version":"1.0","canonical_sha256":"f2cbf463c5dccbe412afff9c7e11ab473d014934a2bb9c1153302556541c61a6","source":{"kind":"arxiv","id":"2404.12387","version":1},"attestation_state":"computed","paper":{"title":"Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CV"],"primary_cat":"cs.CL","authors_text":"Che Zheng, Cyprien de Masson d'Autume, Dani Yogatama, Deyu Fu, Donovan Ong, Eric Chen, Eugenie Lamprecht, Hai Pham, Isaac Ong, Kaloyan Aleksiev, Lei Li, Matthew Henderson, Max Bain, Mikel Artetxe, Nishant Relan, Piotr Padlewski, Qi Liu, Reka Team: Aitor Ormazabal, Ren Chen, Samuel Phua, Yazheng Yang, Yi Tay, Yuqi Wang, Zhihui Xie, Zhongkai Zhu","submitted_at":"2024-04-18T17:59:48Z","abstract_excerpt":"We introduce Reka Core, Flash, and Edge, a series of powerful multimodal language models trained from scratch by Reka. Reka models are able to process and reason with text, images, video, and audio inputs. This technical report discusses details of training some of these models and provides comprehensive evaluation results. We show that Reka Edge and Reka Flash are not only state-of-the-art but also outperform many much larger models, delivering outsized values for their respective compute class. Meanwhile, our most capable and largest model, Reka Core, approaches the best frontier models on b"},"verification_status":{"content_addressed":true,"pith_receipt":true,"author_attested":false,"weak_author_claims":0,"strong_author_claims":0,"externally_anchored":false,"storage_verified":false,"citation_signatures":0,"replication_records":0,"graph_snapshot":true,"references_resolved":false,"formal_links_present":false},"canonical_record":{"source":{"id":"2404.12387","kind":"arxiv","version":1},"metadata":{"license":"http://creativecommons.org/licenses/by/4.0/","primary_cat":"cs.CL","submitted_at":"2024-04-18T17:59:48Z","cross_cats_sorted":["cs.CV"],"title_canon_sha256":"505b6f4604c7286b5200cff3bd555567d1a41abee6b011c49bf6db087f633e63","abstract_canon_sha256":"1f1f198604aff8831b8e1490de480b4dd84d1fa209833904a96376fc2d82aa4d"},"schema_version":"1.0"},"receipt":{"kind":"pith_receipt","key_id":"pith-v1-2026-05","algorithm":"ed25519","signed_at":"2026-07-05T08:16:26.808009Z","signature_b64":"aeXZQENKwijqVq83R7i6xAQEVmXxVIS6SMJ1sFwdCvryZr+e/gJEJbyhCr8AE5UKCOT6vmUY98ypekV8n0z8BA==","signed_message":"canonical_sha256_bytes","builder_version":"pith-number-builder-2026-05-17-v1","receipt_version":"0.3","canonical_sha256":"f2cbf463c5dccbe412afff9c7e11ab473d014934a2bb9c1153302556541c61a6","last_reissued_at":"2026-07-05T08:16:26.807516Z","signature_status":"signed_v1","first_computed_at":"2026-07-05T08:16:26.807516Z","public_key_fingerprint":"8d4b5ee74e4693bcd1df2446408b0d54"},"graph_snapshot":{"paper":{"title":"Reka Core, Flash, and Edge: A Series of Powerful Multimodal Language Models","license":"http://creativecommons.org/licenses/by/4.0/","headline":"","cross_cats":["cs.CV"],"primary_cat":"cs.CL","authors_text":"Che Zheng, Cyprien de Masson d'Autume, Dani Yogatama, Deyu Fu, Donovan Ong, Eric Chen, Eugenie Lamprecht, Hai Pham, Isaac Ong, Kaloyan Aleksiev, Lei Li, Matthew Henderson, Max Bain, Mikel Artetxe, Nishant Relan, Piotr Padlewski, Qi Liu, Reka Team: Aitor Ormazabal, Ren Chen, Samuel Phua, Yazheng Yang, Yi Tay, Yuqi Wang, Zhihui Xie, Zhongkai Zhu","submitted_at":"2024-04-18T17:59:48Z","abstract_excerpt":"We introduce Reka Core, Flash, and Edge, a series of powerful multimodal language models trained from scratch by Reka. Reka models are able to process and reason with text, images, video, and audio inputs. This technical report discusses details of training some of these models and provides comprehensive evaluation results. We show that Reka Edge and Reka Flash are not only state-of-the-art but also outperform many much larger models, delivering outsized values for their respective compute class. Meanwhile, our most capable and largest model, Reka Core, approaches the best frontier models on b"},"claims":{"count":0,"items":[],"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"source":{"id":"2404.12387","kind":"arxiv","version":1},"verdict":{"id":null,"model_set":{},"created_at":null,"strongest_claim":"","one_line_summary":"","pipeline_version":null,"weakest_assumption":"","pith_extraction_headline":""},"integrity":{"clean":true,"summary":{"advisory":0,"critical":0,"by_detector":{},"informational":0},"endpoint":"/pith/2404.12387/integrity.json","findings":[],"available":true,"detectors_run":[],"snapshot_sha256":"c28c3603d3b5d939e8dc4c7e95fa8dfce3d595e45f758748cecf8e644a296938"},"references":{"count":0,"sample":[],"resolved_work":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57","internal_anchors":0},"formal_canon":{"evidence_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"author_claims":{"count":0,"strong_count":0,"snapshot_sha256":"258153158e38e3291e3d48162225fcdb2d5a3ed65a07baac614ab91432fd4f57"},"builder_version":"pith-number-builder-2026-05-17-v1"},"aliases":[{"alias_kind":"arxiv","alias_value":"2404.12387","created_at":"2026-07-05T08:16:26.807575+00:00"},{"alias_kind":"arxiv_version","alias_value":"2404.12387v1","created_at":"2026-07-05T08:16:26.807575+00:00"},{"alias_kind":"doi","alias_value":"10.48550/arxiv.2404.12387","created_at":"2026-07-05T08:16:26.807575+00:00"},{"alias_kind":"pith_short_12","alias_value":"6LF7IY6F3TF6","created_at":"2026-07-05T08:16:26.807575+00:00"},{"alias_kind":"pith_short_16","alias_value":"6LF7IY6F3TF6IEVP","created_at":"2026-07-05T08:16:26.807575+00:00"},{"alias_kind":"pith_short_8","alias_value":"6LF7IY6F","created_at":"2026-07-05T08:16:26.807575+00:00"}],"events":[],"event_summary":{},"paper_claims":[],"inbound_citations":{"count":7,"internal_anchor_count":0,"sample":[{"citing_arxiv_id":"2606.22918","citing_title":"Each Judge Its Own Yardstick: Discovering Per-VLM Taxonomies for Physical Video Evaluation","ref_index":44,"is_internal_anchor":false},{"citing_arxiv_id":"2408.04840","citing_title":"mPLUG-Owl3: Towards Long Image-Sequence Understanding in Multi-Modal Large Language Models","ref_index":175,"is_internal_anchor":false},{"citing_arxiv_id":"2502.04326","citing_title":"WorldSense: Evaluating Real-world Omnimodal Understanding for Multimodal LLMs","ref_index":64,"is_internal_anchor":false},{"citing_arxiv_id":"2602.17555","citing_title":"GraphThinker: Reinforcing Temporally Grounded Video Reasoning with Event Graph Thinking","ref_index":59,"is_internal_anchor":false},{"citing_arxiv_id":"2404.16821","citing_title":"How Far Are We to GPT-4V? Closing the Gap to Commercial Multimodal Models with Open-Source Suites","ref_index":89,"is_internal_anchor":false},{"citing_arxiv_id":"2406.16852","citing_title":"Long Context Transfer from Language to Vision","ref_index":57,"is_internal_anchor":false},{"citing_arxiv_id":"2406.07476","citing_title":"VideoLLaMA 2: Advancing Spatial-Temporal Modeling and Audio Understanding in Video-LLMs","ref_index":39,"is_internal_anchor":false}]},"formal_canon":{"evidence_count":0,"sample":[],"anchors":[]},"links":{"html":"https://pith.science/pith/6LF7IY6F3TF6IEVP76OH4ENLI4","json":"https://pith.science/pith/6LF7IY6F3TF6IEVP76OH4ENLI4.json","graph_json":"https://pith.science/api/pith-number/6LF7IY6F3TF6IEVP76OH4ENLI4/graph.json","events_json":"https://pith.science/api/pith-number/6LF7IY6F3TF6IEVP76OH4ENLI4/events.json","paper":"https://pith.science/paper/6LF7IY6F"},"agent_actions":{"view_html":"https://pith.science/pith/6LF7IY6F3TF6IEVP76OH4ENLI4","download_json":"https://pith.science/pith/6LF7IY6F3TF6IEVP76OH4ENLI4.json","view_paper":"https://pith.science/paper/6LF7IY6F","resolve_alias":"https://pith.science/api/pith-number/resolve?arxiv=2404.12387&json=true","fetch_graph":"https://pith.science/api/pith-number/6LF7IY6F3TF6IEVP76OH4ENLI4/graph.json","fetch_events":"https://pith.science/api/pith-number/6LF7IY6F3TF6IEVP76OH4ENLI4/events.json","actions":{"anchor_timestamp":"https://pith.science/pith/6LF7IY6F3TF6IEVP76OH4ENLI4/action/timestamp_anchor","attest_storage":"https://pith.science/pith/6LF7IY6F3TF6IEVP76OH4ENLI4/action/storage_attestation","attest_author":"https://pith.science/pith/6LF7IY6F3TF6IEVP76OH4ENLI4/action/author_attestation","sign_citation":"https://pith.science/pith/6LF7IY6F3TF6IEVP76OH4ENLI4/action/citation_signature","submit_replication":"https://pith.science/pith/6LF7IY6F3TF6IEVP76OH4ENLI4/action/replication_record"}},"created_at":"2026-07-05T08:16:26.807575+00:00","updated_at":"2026-07-05T08:16:26.807575+00:00"}