Pith. sign in

REVIEW 3 cited by

ACORN: Adaptive Coordinate Networks for Neural Scene Representation

Not yet reviewed by Pith; the record is open.

This paper has not been read by Pith yet. Machine review is queued; the pith claim, tier, and objections will appear here once it completes.

SPECIMEN: schema-true, not a live event

T0 review · schema-true

One-sentence machine reading of the paper's core claim.

pith:XXXXXXXX · record.json · timestamp

arxiv 2105.02788 v1 pith:X2LAPU3W submitted 2021-05-06 cs.CV cs.GRcs.LG

classification cs.CVcs.GRcs.LG
keywords networkneuralrepresentationstrainingarchitectureimagesrepresentapproach
verification ladder T0 review T1 audit T2 compute T3 formal
0 comments
read the original abstract

Neural representations have emerged as a new paradigm for applications in rendering, imaging, geometric modeling, and simulation. Compared to traditional representations such as meshes, point clouds, or volumes they can be flexibly incorporated into differentiable learning-based pipelines. While recent improvements to neural representations now make it possible to represent signals with fine details at moderate resolutions (e.g., for images and 3D shapes), adequately representing large-scale or complex scenes has proven a challenge. Current neural representations fail to accurately represent images at resolutions greater than a megapixel or 3D scenes with more than a few hundred thousand polygons. Here, we introduce a new hybrid implicit-explicit network architecture and training strategy that adaptively allocates resources during training and inference based on the local complexity of a signal of interest. Our approach uses a multiscale block-coordinate decomposition, similar to a quadtree or octree, that is optimized during training. The network architecture operates in two stages: using the bulk of the network parameters, a coordinate encoder generates a feature grid in a single forward pass. Then, hundreds or thousands of samples within each block can be efficiently evaluated using a lightweight feature decoder. With this hybrid implicit-explicit network architecture, we demonstrate the first experiments that fit gigapixel images to nearly 40 dB peak signal-to-noise ratio. Notably this represents an increase in scale of over 1000x compared to the resolution of previously demonstrated image-fitting experiments. Moreover, our approach is able to represent 3D shapes significantly faster and better than previous techniques; it reduces training times from days to hours or minutes and memory requirements by over an order of magnitude.

Discussion (0). Continue with ORCID to comment.

Forward citations

Cited by 3 Pith papers

Reviewed papers in the Pith corpus that reference this work. Sorted by Pith novelty score. Full citation record

  1. Instant GaussianImage: A Generalizable and Self-Adaptive Image Representation via 2D Gaussian Splatting

    cs.CV 2025-06 conditional novelty 6.0 of 10

    A learnable initialization network plus short fine-tuning produces 2D Gaussian image representations faster than GaussianImage, with adaptive Gaussian counts per image.

  2. NAT: Neural Acoustic Transfer for Interactive Scenes in Real Time

    cs.SD 2025-06 conditional novelty 6.0 of 10

    NAT predicts acoustic transfer maps for dynamically changing scenes in 1-4 ms using a neural field trained on fast boundary element simulations.

  3. Large Images are Gaussians: High-Quality Large Image Representation with Levels of 2D Gaussian Splatting

    cs.CV 2025-02 conditional novelty 6.0 of 10

    A two-level 2D Gaussian splatting method with direct covariance optimization fits large images with more Gaussian points and higher PSNR than prior Gaussian-based image representation.

Pith tools