Pith. sign in

Paper Citation Record · LEDGER

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving

As of 21 August 2026, this Paper Citation Record lists 25 of 25 outbound references and 1 inbound Pith citation observation for arXiv:2507.00707.

A citation records a reference. It does not transfer a finding from one paper to another.

pith.paper-citation-record.v1
2507.00707 v1

Coverage vector

measured 25 of 25 reference resolution

Typed states for the displayed outbound observations.

Source: paper_references, paper_reference_links, observed 2026-08-06T21:16:03.985223Z

measured 26 of 26 standing notices

One-hop event checks from named stored sources.

Source: scholarly_work_events, retraction_status_cache, observed 2026-08-21T06:32:19.484+00:00

measured 1 of 1 inbound itemization

Pith citing papers itemized under the disclosed page cap.

Source: paper_references, paper_reference_links, observed 2026-08-02T01:23:06.577719Z

measured 0 of 1 external citation measurements

A source-named dated measurement, never combined with another source.

Source: cited_works

Reference resolution

25 of 25 outbound references displayed

  • verified exact0
  • verified fuzzy17
  • unresolved8
  • parse uncertain0
  • malformed identifier0
  • metadata mismatch0

External citation measurements

No source-named external measurement is stored.

Outbound references

Observation 769ed58f-a6e8-4c27-a985-0697b7262a43 · outbound

This paper cites MagicDrive: Street View Generation with Diverse 3D Geometry Control.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving MagicDrive: Street View Generation with Diverse 3D Geometry Control

Reference 1

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:02.509759Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:02.509759Z digest=sha256:435a1c89195b0da108d63ea2b989bf89284361cedf5f5fbc6f5a7b8077d7535f

Observation 2d42f68f-3106-496a-8c2a-f72eabd645e4 · outbound

This paper cites Drivingdiffusion: Layout-guided multi-view driving scenarios video generation with latent diffusion model.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Drivingdiffusion: Layout-guided multi-view driving scenarios video generation with latent diffusion model

Reference 2

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.591249Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:02.561112Z digest=sha256:45fcb35c4ac3a18f54b4eb9068ffc7146dcab328166e55b23a71f2145fced642

Observation a08a8240-46a5-4bf8-a291-8ee4fb7feb20 · outbound

This paper cites Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Driving into the future: Multiview visual forecasting and planning with world model for autonomous driving

Reference 3

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.496258Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:02.664596Z digest=sha256:1ffd4f5a383f65366c1fbe31a16cf01cc9025cd020a413c2d10ae4c61ab49af3

Observation 0aea0997-6f15-4b82-8a56-087ebb3d8108 · outbound

This paper cites Panacea: Panoramic and controllable video generation for autonomous driving.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Panacea: Panoramic and controllable video generation for autonomous driving

Reference 4

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.349411Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:02.708374Z digest=sha256:8d0487f41503698c92d5acc79411b28dfa37e925d2ece2b304b6e5ef8d54485d

Observation dedffad0-b373-4094-84be-9228e6e88250 · outbound

This paper cites Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Lift, splat, shoot: Encoding images from arbitrary camera rigs by implicitly unprojecting to 3d

Reference 5

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.220621Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:02.788028Z digest=sha256:da75f97f4b9c6483e8c5c69e0a3eb21d8a71cbcac393e38c3a4554f17c73c595

Observation 20d6d946-f60b-4246-b128-04f842f7e7e0 · outbound

This paper cites BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving BEVDet: High-performance Multi-camera 3D Object Detection in Bird-Eye-View

Reference 6

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:02.821152Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:02.821152Z digest=sha256:e041a4194e415b893f12be63c06ff6f115d8b89cc8f2debd29ff23b0ceecd58f

Observation bc3292b6-0aea-43c0-821c-020df08a96bb · outbound

This paper cites Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Bevfusion: Multi-task multi-sensor fusion with unified bird’s-eye view representation

Reference 7

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:06.095644Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:02.876874Z digest=sha256:6ba2cd010d70cfb21daec72707b37be33e30178e6e727e4885d1295410db23c4

Observation 7d8bc499-4011-492a-8171-59a739b5ecd2 · outbound

This paper cites Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal trans- formers.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Bevformer: Learning bird’s-eye-view representation from multi-camera images via spatiotemporal trans- formers

Reference 8

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.932218Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:02.916508Z digest=sha256:54c9f8a7577d16533259c686b722d9ca5daa3c49e6c305b9f06b5d5e8c3f1e3d

Observation b1327940-5b5d-456f-990f-f3ab4fa2f54a · outbound

This paper cites Planning-oriented autonomous driving.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Planning-oriented autonomous driving

Reference 9

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.864956Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:02.947076Z digest=sha256:5e80d55ccb73c882c0a28647366842b468344af9179a6a85209fe0a1ce23d2f0

Observation 41cdfdbb-c5af-4ee5-aa2d-94df67354be3 · outbound

This paper cites Neural discrete representation learning.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Neural discrete representation learning

Reference 10

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:02.972118Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:02.972118Z digest=sha256:d6c5b0ee0aafa63a56d7faa3204198e89c51e7ced5a073a1ced97e10b8d67f51

Observation 2257872d-bf9c-41fd-921e-c46278b9be47 · outbound

This paper cites Taming transformers for high-resolution image synthesis.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Taming transformers for high-resolution image synthesis

Reference 11

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.708350Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.040381Z digest=sha256:d3a02580f2e0465beed0c370d00c85234954c5b09c4fcedb0fe8b2ef9572a1cd

Observation bad47d59-1fd0-4b00-ac7c-46dee8ec2bec · outbound

This paper cites Generative adversarial nets.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Generative adversarial nets

Reference 12

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.085635Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.085635Z digest=sha256:1a04d1eb86a8839d72e46115b349aec93d5f3668853aa925fd60bad2262941eb

Observation 4d9b0286-5eee-429b-bbc6-ffa31108cd44 · outbound

This paper cites Perceptual losses for real-time style transfer and super- resolution.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Perceptual losses for real-time style transfer and super- resolution

Reference 13

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.567907Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.122390Z digest=sha256:e78b228933d9343002d33e26f662002255d67b3f994ef8a10369aac953663706

Observation 4292a0d8-a245-4237-9f74-c6be65f04009 · outbound

This paper cites Vector-quantized Image Modeling with Improved VQGAN.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Vector-quantized Image Modeling with Improved VQGAN

Reference 14

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.158583Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.158583Z digest=sha256:3660fc0cd2c8b5b314221d81f250528f55ab602cd58026dc8e7415b61a43aacd

Observation cf18d2c8-c8e9-4044-8b52-e7fa5819b09d · outbound

This paper cites An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving An Image is Worth 16x16 Words: Transformers for Image Recognition at Scale

Reference 15

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.199718Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.199718Z digest=sha256:46579b6318f366fd853fa96d05395db29fe3751485544b3e80e7fb2bba1f95c7

Observation a4cd8744-61ce-4a54-ad7d-f2f08314b127 · outbound

This paper cites Denoising diffusion probabilistic models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Denoising diffusion probabilistic models

Reference 16

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.249349Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.249349Z digest=sha256:9cb361c88f713c359788d5b0e1847087406ef78a93016050f5d365af3056f212

Observation 70a4ed66-17e0-4a35-b601-b46b0efe62ab · outbound

This paper cites Denoising Diffusion Implicit Models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Denoising Diffusion Implicit Models

Reference 17

Resolution
unresolved
no resolver link, observed 2026-08-06T21:16:03.347096Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-06T21:16:03.347096Z digest=sha256:04494b57a9d123f40dc4eedb25234125b1b8dd0e122db52fc95d97b739ed4164

Observation 461e5eb7-26cf-4b27-97d6-155d1422d189 · outbound

This paper cites High-resolution image synthesis with latent diffusion models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving High-resolution image synthesis with latent diffusion models

Reference 18

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.434381Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.382308Z digest=sha256:6aed06d8b3473bddcc23bcbef598ef50e10784dccad04ef529347154b4f0b781

Observation 0b771eb2-ae24-431b-ba13-cdd9712f509e · outbound

This paper cites Scalable diffusion models with transformers.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Scalable diffusion models with transformers

Reference 19

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.262604Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.466342Z digest=sha256:6346548678f0ffeeab640e98685dd0c5cda8ddeb10ad8294fc3b75ba7a0b5755

Observation b028a2df-cc99-4301-b53a-73e3d895b032 · outbound

This paper cites Street-view image generation from a bird’s-eye view layout.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Street-view image generation from a bird’s-eye view layout

Reference 20

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.169444Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.574745Z digest=sha256:7b4d73eecd3f0f671bd30c324c7045fa340f913e039846f1ac5461ca01f0de27

Observation 6d9dfdbf-2b59-41d4-8021-5289a736b6f4 · outbound

This paper cites Nerf: Representing scenes as neural radiance fields for view synthesis.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Nerf: Representing scenes as neural radiance fields for view synthesis

Reference 21

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:05.063741Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.662211Z digest=sha256:3c592caa0c01668442d1f8f921df2589ebc5f5c0118a743ac232c1f3a234caf2

Observation e4caaf21-99d8-4422-bdf4-d9893319697d · outbound

This paper cites Adding conditional control to text-to-image diffusion models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Adding conditional control to text-to-image diffusion models

Reference 22

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:04.897313Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.734048Z digest=sha256:6b46b23fafa600a1003ba981ecad5d2ef1e9ecfd9c88a2a2aae43f26dad31d6c

Observation b1852ddb-eacc-40e9-9029-fe74ada547e5 · outbound

This paper cites Feature pyramid networks for object detection.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Feature pyramid networks for object detection

Reference 23

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:04.679187Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.802897Z digest=sha256:3beb3caf0b670ca6ffd3f76869571d9d34cafd2ede0e0690bd6f50e47a9162e0

Observation dd832207-7098-46cb-aa28-d4769f758fd6 · outbound

This paper cites Loftr: Detector-free local feature matching with transformers.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Loftr: Detector-free local feature matching with transformers

Reference 24

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:04.447945Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.905488Z digest=sha256:2833382a817c0ba86920e8d0b214aefe65e923e131d80d05d418e4fa207281d9

Observation f846a9e7-da1a-44c9-9d91-44a6ce011637 · outbound

This paper cites Laion-5b: An open large-scale dataset for training next generation image-text models.

BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving Laion-5b: An open large-scale dataset for training next generation image-text models

Reference 25

Resolution
verified fuzzy
raw_fallback, observed 2026-08-06T21:16:04.213139Z

Source-reported events for the cited work

No event found in the named queried sources as of 2026-08-21T06:32:19.484+00:00.

source=pdf_text observed=2026-08-06T21:16:03.985223Z digest=sha256:a832e24d3fc7cb8c5f50cf1103cb4904b32c28d72aed3c228507d189f1eedeb7

Pith citing papers

Observation dcea4640-d784-4eb3-bbef-fedbc86f5832 · inbound

Variational Inference for Bird's Eye View Segmentation in Autonomous Driving cites this paper.

Variational Inference for Bird's Eye View Segmentation in Autonomous Driving BEV-VAE: Multi-view Image Generation with Spatial Consistency for Autonomous Driving

Reference 44

Resolution
unresolved
no resolver link, observed 2026-08-02T01:23:06.577719Z

Source-reported events for the cited work

Unavailable: canonical work link unavailable.

source=pdf_text observed=2026-08-02T01:23:06.577719Z digest=sha256:3d46d05c22b60940774c4503a23c5d2a6dd1177931acff884f210584ceac70ed