CHAPTER 5. UNCERTAINTY-AWARE NEURAL NETWORKS FOR HIGH RESOLUTION DIGITIZATION

Figure 5.1: From a single image captured with a flatbed scanner UMat estimates high-resolution SVBRDFs. Leveraging an attention-guided generative model, the estimations are sharp and artifact-free, and can be used in photorealistic rendering. We introduce a novel uncertainty quantification algorithm for material digitization, which correlates with error at test time and can be used for dataset creation.

Figure 5.1: From a single image captured with a flatbed scanner UMat estimates high-resolution SVBRDFs. Leveraging an attention-guided generative model, the estimations are sharp and artifact-free, and can be used in photorealistic rendering. We introduce a novel uncertainty quantification algorithm for material digitization, which correlates with error at test time and can be used for dataset creation.

In this chapter, we introduce a learning-based method to recover normals, specularity, and roughness from a single diffuse image of a material, using microgeometry appearance as our primary cue. Previous methods that work on single images tend to produce over-smooth outputs with artifacts, operate at limited resolution, or train one model per class with little room for generalization. In contrast, in this work, we propose a novel capture approach that leverages a generative network with attention and a U-Net discriminator, which shows outstanding performance integrating global information at reduced computational complexity. We showcase the performance of our method with a real dataset of digitized textile materials and show that a commodity flatbed scanner can produce the type of diffuse illumination required as input to our method. Additionally, because the problem might be ill-posed śmore than a single diffuse image might be needed to disambiguate the specular reflectionś or because the training dataset is not representative enough of the real distribution, we propose a novel framework to quantify the model’s confidence about its predictions at test time. Our method is the first one to deal with the problem of modeling uncertainty in material digitization, increasing the trustworthiness of the process and enabling more intelligent strategies for dataset creation, as we demonstrate with an active learning experiment. A version of the model presented in this chapter is part of Textura.ai. The contributions in this chapter led to the following publication1:

“UMat: Uncertainty-Aware Single Image High Resolution Material Capture”

Carlos Rodriguez-Pardo, Henar Dominguez, David Pascual, Elena Garces

Proc. of Computer Vision and Pattern Recognition (CVPR) (2023)

Figure 5.2: Our method digitizes a material taking as input a single scanned image. Further, it returns a pixel-wise metric of uncertainty σBRDF, computed at test time through probabilistic sampling, proven useful for active learning. In the plot we compare the average deviations of the radiance of different renders in the blue crop w.r.t the ground truth (GT) of: 1) the distribution of the probabilistic samples of a model trained with 100% of the data; 2) the deterministic output of that model; 3) the output of a model trained using 40% of the training dataset, sampled by active learning guided by σBRDF and; 4) a model trained using 40% of the training dataset, randomly sampled. The material at the bottom, for which the model shows a higher uncertainty, generates more varied renders and differs most from the ground truth.

Figure 5.2: Our method digitizes a material taking as input a single scanned image. Further, it returns a pixel-wise metric of uncertainty σBRDF, computed at test time through probabilistic sampling, proven useful for active learning. In the plot we compare the average deviations of the radiance of different renders in the blue crop w.r.t the ground truth (GT) of: 1) the distribution of the probabilistic samples of a model trained with 100% of the data; 2) the deterministic output of that model; 3) the output of a model trained using 40% of the training dataset, sampled by active learning guided by σBRDF and; 4) a model trained using 40% of the training dataset, randomly sampled. The material at the bottom, for which the model shows a higher uncertainty, generates more varied renders and differs most from the ground truth.

5.1 Introduction

Virtual design, online marketplaces, product lifecycle workflows, AR/VR, videogames, . . . , all require lifelike digital representations of real-world materials (i.e., digital twins). Acquiring these digital copies is typically a cumbersome and slow process that requires expensive machines and several manual steps, creating roadblocks for scalability, repeatability, and consistency. Among the many industries requiring digital twins of materials, the fashion industry is in a critical position; facing the demand to digitize hundreds of samples of textiles in short periods, which cannot be achieved with the current technology.

In this context, casual capture systems for optical digitization provide a promising path for scalability. These systems leverage handheld devices (such as smartphones), one or more different illuminations, and learning-based priors to estimate the material’s diffuse and specular reflection lobes. However, existing approaches present several drawbacks that

make them unsuitable for practical digitization workflows. Generative solutions [Guo+20b; VPS21] typically produce artifacts and are generally not tileable. Despite recent attempts to improve tileability and controllability [Zho+22], these solutions are slow to train and evaluate (requiring online optimization iterations), are limited in resolution, and present challenges for generalization (requiring one model per material class). Further, the fact that these methods build on perceptual losses śnot pixel lossesś to compare the input photo with the generated material entails extra difficulties when it comes to guaranteeing the repeatability and consistency required for building a digital inventory (i.e., color swatches, prints, or other variations). On the other hand, methods that build on differentiable node graphs [Hen+21] overcome the tileability and resolution limitations, yet, they share the problems derived from using perceptual losses and category-specific training.

In this chapter, we present UMat, a practical, scalable, and reliable approach to digitizing the optical appearance of textile material samples using SVBRDFs. Commonly used SVBRDFs typically contain two reflection terms: a diffuse term, parameterized by an albedo image, and a specular one, parameterized by normals, specularity, and roughness. Prior work typically estimates both components, which becomes a very challenging problem, obtaining over-smooth outputs, and being prone to artifacts [Des+18; Guo+21; ZK21]. Instead, in this chapter, we demonstrate that it is possible to provide accurate digitizations of materials leveraging as input a single diffuse image that acts as albedo and estimating the specular components using a neural network.

Our key observation is to realize that most of the appearance variability of textile materials is due to its microgeometry and that a commodity flatbed scanner can approximate the type of diffuse illumination that we require for the majority of textile materials (see Figure 5.3).

Figure 5.3: Scanner images vs fitted albedos.

Figure 5.3: Scanner images vs fitted albedos.

Nevertheless, single-image material estimation is still an ill-posed problem in our setting, as reflectance properties may not be directly observable from a single diffuse image. To account for these non-directly observable properties, we propose a novel way to measure the model’s confidence about its prediction at test time. Leveraging Monte Carlo (MC) Dropout [GG16], we propose an uncertainty metric computed as the variance of sampling and evaluating multiple estimations for a single input in a render space. We show that this confidence directly correlates with the accuracy of the digitization, which helps identify ambiguous inputs, out-of-distribution samples, or under-represented classes. Besides increasing the trustworthiness of the capture process, our confidence quantification enables smarter strategies for dataset creation, as we demonstrate with an active learning experiment.

We pose the estimation as an Image-to-Image Translation problem (I2IT) that directly re-gresses roughness, specular, and normals, from a single input image. Under the hood, our novel residual architecture has a single encoder enhanced with lightweight attention modules [Wan+20b; MR21] for improving global consistency and reducing artifacts, specialized decoders for each target reflectance map, and a U-Net discriminator [SSK20], which enhances generalization.

In summary, we present the following contributions:

• A novel material capture system which leverages the diffuse illumination provided by flatbed scanners for high-resolution, scalable, and reliable digitizations.

• An attention-enhanced GAN model and training procedure designed for maximizing accuracy and sharpness, and removing undesired artifacts.

• A generic uncertainty quantification framework for material capture algorithms which correlates with prediction error on a render space.

• An exhaustive evaluation method for measuring the accuracy and quality of the model estimations.

5.2 Related Work

5.2.1 Lightweight Material Capture

Single Image SVBRDF Capture Capturing accurate SVBRDFs from a single image is a challenging problem that requires predicting the photometric response of a material given only a sample of it. The most common approximation is to use a flash-lit front planar image captured with a smartphone. Extending neural style transfer [GEB15b] to material capture, Aittala et al. [AAL16] leverage pre-trained CNNs and texture priors for smartphone material acquisition. Relatedly, Henzler et al. [Hen+21] use style losses for training generative models for BRDF synthesis. These approaches are optimized for synthesis, which allow for seamlessly tileable outputs, and do not require supervised training. However, they work best for stochastic materials, limiting their scope.

Leveraging datasets of labeled materials, different methods have trained autoencoders for SVBRDF capture. Li et al. [Li+17a] reconstruct spatially-varying albedo and normals using a U-Net [RFB15], and homogeneous specular albedo and roughness using a CNN regressor. This work was extended through Self-Augmented CNNs in [Ye+18]. By leveraging Conditional Random Fields (CRFs), a material classifier and one decoder per map, Li et al. [LSC18] reconstruct spatially-varying albedo, roughness, and normals. Deschaintre et al. [Des+18] propose a modified U-Net, synthetic datasets, and a render loss. Cascaded models [Li+18; SC20]; and deep latent spaces optimized using inverse rendering [Gao+19] have also shown success for this problem. Recently, Generative Adversarial Networks (GANs) have shown improved capabilities compared to more naive losses. These require a discriminator, which can be trained on renders [Wen+22], SVBRDF maps [Guo+21; VPS21], or both [ZK21].

Our approach differs from previous work in several factors. Importantly, we use flatbed scanners instead of smartphones for material capture. While they limit the materials which can be captured, they provide adequate illumination for easier digitizations, and a higher level of resolution and detail. We hypothesize that material specularity can be estimated accurately by leveraging its microgeometry. From this assumption, we build a GAN which, in contrast with previous work, leverages state-of-the-art attention mechanisms and discriminator design for obtaining a more holistic understanding of its inputs. Further, we train exclusively on real data, and propose a more comprehensive evaluation.

Multiple Image SVBRDF Capture A different corpus of work relies on multiple images of the material. By capturing more evidence of the material photometric response, they provide more accurate SVBRDFs. This hinders scalability, as they require a larger capture and calibration effort. The simplest approach is to require two samples, a diffuse and a flash-lit image [AWL15; Bos+20]. More flexible alternatives allow for more samples, combined with a learned prior [Guo+20b; Des+19]; or Monte Carlo rendering [Lua+21]. A different approach is to capture the material at a high resolution, and transfer those details to larger samples of the same material [DDB20; RG21]. Finally, videos for material acquisition have also shown to be promising capture systems [Ye+21].

Procedural Graphs Procedural graphs have also been used for material generation. Instead of relying on material priors, these approaches work by optimizing a material graph through a differentiable pipeline. These provide interesting capabilities, such as easy edition or tiling, but are limited by the expressiveness of the procedural model. These have been explored for general SVBRDF estimation [Shi+20; Hu+22c; Guo+20a], or for high-quality woven fabric digitizations [Jin+22].

5.2.2 Uncertainty Quantification in Deep Learning

Measuring the confidence of deep learning models is an active research area [KG17] with multiple applications, including safety-critical problems, like self-driving [Hua+19] or medicine [Kur+22]; and active dataset creation [Lei+21; Sol+21]. In computer vision, uncertainty quantification

has focused on image classification [SKK18; Lei+21], segmentation [Kry+21], and depth regression [Hu+22a; Ami+20]. To overcome the computational intractability of Bayesian Neural Networks, different approximations have been proposed, including MC Dropout [GG16], Deep Ensembles [LPB17] or Variational Inference [MNG17]. Orthogonal alternatives exist, including Evidential Deep Learning [Ami+20; SKK18] or frequentist approaches like Constrained Ordinal Regression [Hu+22a]. We refer the reader to recent surveys [Jos+22; Gaw+21] for more comprehensive reviews.

Quantifying uncertainties allows communicating the end user that the model predictions may be inaccurate, suggesting alternative pathways; as well as cheaper dataset creation through active learning. Bayesian material parameter estimation has been proposed for procedural frameworks [Guo+20a], but, to the best of our knowledge, it has not been explored for SVBRDF estimation. We aim to propose an efficient uncertainty quantification framework for deep SVBRDF capture methods which accounts for material perception and is agnostic to the material model.

5.3 Method

Our method starts from an input image X of a material taken under diffuse lighting and outputs the parameters of the specular lobe of the SVBRDF, i.e., M={Mi}i=13 corresponding to the material roughness, specularity, and normals. We illustrate this process in Figure 5.4.

Figure 5.4: Overview of UMat. We propose an attention-guided generator trained with style, frequency, pixel-wise, and adversarial losses. In green, we show the components that include any form of attention mechanism. The supplementary material contains the detailed architectures. On the right, we show two applications of our method: First, thanks to our test-time uncertainty quantification, we can provide a measure of the reliability of the estimation. Second, the maps that UMat produces can be used by any render engine.

Figure 5.4: Overview of UMat. We propose an attention-guided generator trained with style, frequency, pixel-wise, and adversarial losses. In green, we show the components that include any form of attention mechanism. The supplementary material contains the detailed architectures. On the right, we show two applications of our method: First, thanks to our test-time uncertainty quantification, we can provide a measure of the reliability of the estimation. Second, the maps that UMat produces can be used by any render engine.

Following previous work [KG13; Mun+22], we use the physically-based material model from Disney [BS12], which aggregates a diffuse term with an isotropic, microfacet specular GGX lobe s(M) [Wal+07], such that, fl,v(M,X)=Xπ+sl,v(M)x×y is the shading model for a light l and view position v. We formulate this estimation as an Image-to-Image Translation problem (I2IT). We train a U-Net [RFB15] generator G(X)=M^ within a GAN framework, extending the adversarial loss with pixel-wise, style, and frequency losses. We design the generator to maximize accuracy, sharpness, and robustness. To do so, we train a single encoder, enhanced with self-attention layers and a transformer, and use one decoder per map. Sections 5.3.1, 5.3.2, and 5.3.3 present the network design, the loss and data augmentation choices, respectively. Implementation details are provided on the supplementary material (Section 5.A).

Using as input a single image taken under diffuse lighting presents extra challenges when estimating the SVBRDF; we lack the extra cues provided by more complex illumination patterns (e.g. flash lighting). Therefore, in Section 5.4, we explain how to compensate for this potential ambiguity by introducing an uncertainty metric that can be computed at test time. Section 5.5 presents the evaluation, which includes the description of our dataset and metrics (Sections 5.5.1 and 5.5.2), an ablation study that validates the design (5.6.1), qualitative and quantitative results (5.6.2 and 5.6.3), an application of our uncertainty metric in an active learning setting (5.6.4), and comparisons with previous work (5.6.5).

5.3.1 Network Design

We use a U-Net [RFB15] with residual connections [He+16; Dia+20; DLG21] in all of our convolutional blocks, an individual decoder per map [Gar+22; RG22; DLG21; ZK21], and group normalization [WH18]. To each decoder, we add a pixel-wise dropout-regularized MLP [Sri+14], aimed at increasing the accuracy of the predictions and allowing us to measure uncertainty at test-time.

While this multi-decoder residual U-Net is relatively accurate, it is limited by its receptive field, as is common on fully-convolutional architectures. Previous work [Des+18] proposed the use of a global track for fusing spatially distant information. We instead draw inspiration from recent advances in attention and diffusion models [Sah+22; HJA20; Rom+22], and add a self-attention module with linear complexity [Wan+20b] to the output of every convolutional block in the encoder. Finally, we add a lightweight MobileViT [MR21] to the bottleneck to provide the model with a global understanding of its input. By performing the most complex computations at the encoder, we provide the specialized decoders with dense inputs which are computed only once.

5.3.2 Loss Function

LG=iλiLpixeli+λadvLadv+λstyleLstyle+λfreqLfreq                  (5.1)

Lpixel is the L1 norm weighted per map, λi. L1 produces sharper results than higher-order alternatives, such as L2. We introduce an adversarial loss to handle the intrinsic ambiguity of ill-posed problems [Tex+20b; Iso+17; Gar+22; VPS21]. In our case, the choice of the discriminator is a critical design decision.

Recent work [SSK20] proposed U-Net architectures for discriminators, which allows for better learn both low and high-level features, and introduces further regularization. These result in more conservative albeit less diverse generations [Han+22a]. We use a U-Net discriminator [SSK20] with attention [Woo+18], which outputs two estimations: a scalar output Denc, provided by its encoder, and a 2D estimation Ddec, provided by its decoder. Denc provides a global estimate of the quality of the stack M, while Ddec gives pi estimations. As in [SSK20], we add a regularization term LDdeccons, and leverage cut-mix as for discriminator data-augmentation. Our discriminator and adversarial losses are:

LD=LDenc+LDdec+λconsLDdeccons                  (5.2)

Ladv=log(Denc(G(X))+log(Ddec(G(X))                  (5.3)

where the implementation of LDenc and LDdec follows [SSK20].

We further add two losses to improve the accuracy and sharpness of the results: a frequency loss Lfreq and a style loss Lstyle. Lfreq is estimated by averaging the focal frequency loss [Jia+21] computed over each individual channel of M. This is designed to help GANs preserve high-frequency details. Further, as shown in prior work, working in the frequency domain is beneficial when handling textures [AWL13; Mar+20]. Our style loss Lstyle is inspired by the success of neural losses when dealing with textures [GEB15b; GEB15a]. However, off-the-shelf metrics which are designed for 3-channel images, are not immediately usable in SVBRDF. While it is possible to compute them for each map separetely [RG22], this does not necessarily preserve inter-map consistency. We follow recent work [CHB21] and use LPIPS [Zha+18] taking as input a 3-channel image created by randomly sampling three channels from the set of five available channels of M.

5.3.3 Data Augmentation

We follow two strategies for data augmentation. First, we perform patch-based training and affine transforms with randomly cropped patches [RG21; Tex+20b; VPS21]. We also apply random rescales for generalization at lower resolutions, and rotations to account for possible misalignments that may arise when capturing the samples. Second, we apply several image transformations to increase model robustness: random intensity changes in HSV space to the inputs, gaussian noise and blurs, and random erasing [Zho+20] for regularization.

5.4 Uncertainty Quantification

Material acquisition from a single diffuse image is potentially an ill-posed problem, since it assumes that the microgeometry is a sufficient cue to predict the material appearance. While purely learned priors have been proven to work well for inverse problems of single images object shape estimation [Lic+21; Lua+21; Hwa+22; Mun+22; Li+18], they are typically combined with render losses that guarantee consistency in the reconstruction. We lack the necessary input to include this kind of supervision, therefore, we propose an alternative approach aimed at quantifying the confidence of the prediction.

This is valuable for several purposes. First, it provides a way of communicating possible inaccuracies to users of these systems when it is not possible to have access to the ground truth reflectance values; and, importantly, it enables efficient dataset creation through active learning, as we show in Section 5.6.4.

We propose an uncertainty quantification mechanism that is applied to individual per-map estimations, and also globally in a render space.

It is possible to measure uncertainty, among other methods, through deep ensembles [LPB17], evidential learning [SKK18; BYK21; Wan+22b], or ordinal regression [Hu+22a]. However, they are costly to train and evaluate, and would imply major changes in our method. Instead, we follow a probabilistic approach called MC Dropout [GG16], with which, for a particular input, we sample a set of predictions by adding randomness to the forward pass of our model. This process has no impact in the regular deterministic evaluation and implies no changes in our model architecture. Specifically, while measuring uncertainty, we randomly deactivate 20% of the neurons of the MLP of our decoders, obtaining a set of outputs U={M^j~G(X)}j=1N that we use to compute several metrics.

First, we compute the pixel-wise standard deviation for each invididual map separately, obtaining σ<),σspec, and σrough, for the normals, specular, and roughness maps, respectively. Then, inspired by perceptual metrics for computing BRDF differences [Lav+21], we define our novel perceptually-aware uncertainty metric σBRDF as follows:

σBRDF=1|xy|xylog(1|S|(l,v)Sσl,v2({fl,v(Uj,K)cos(θl)}j=1N)3)                  (5.4)

where K is a 2D grayscale image of a constant neutral grey value, σl,vx×y is the pixel-wise standard deviation of the renders obtained for the set of sampled maps U, and S is the fixed set of 50 views optimized in [NJR15] for efficient BRDF capture. The log spatially-varying result is integrated across the spatial dimensions xy to obtain a single value as output. This equation introduces a perceptual component to our uncertainty metric in two forms: first, by applying cosine weighting to the light position to compensate for light attenuation at grazing angles, and second, by taking the cubic root of these differences to attenuate peak reflectances. Figure 5.5 showcases an example of the predicted uncertainties for a rib material with sequins. While the uncertainty is low at the yarns because it is a common material in our dataset, it appears high in the sequins since that effect had not been observed during training.

Figure 5.5: Top: input image of a rib material with metallic sequins. Bottom: σBRDF and per-map uncertainties.

Figure 5.5: Top: input image of a rib material with metallic sequins. Bottom: σBRDF and per-map uncertainties.

5.5 Experimental Results

5.5.1 Dataset

We gathered a novel dataset for training and testing our method. It comprises 2000 textile materials of a variety of families and microstructures that we divided into 14 families: crepe, jacquard, pile, plain, satin, and twill for wovens; fleece, french terry, interlock, jersey, milano, pique, and rib for knits; and, finally, leathers. Each family differs in its construction pattern

(i.e., its microstructure), which directly impacts its optical appearance. In the supplementary material, we show a more detailed analysis of the dataset. For each material, we have an image scanned with the flatbed scanner EPSON V850 whose lighting configuration is close to diffuse (as we show in Figure 5.3), and its corresponding ground truth specular maps of the SVBRDF (normals, specular, and roughness). To obtain these maps, first, we digitized the material with an optical gonioreflectometer and then, propagated the maps to the scan using map propagation techniques [RG21]. All our images and maps have a resolution of 1000 PPIs, allowing us to leverage the full semantics of the microstructure for inference. We split our dataset into 90-10 for train and test, making sure that every family is equally represented in both splits.

5.5.2 Metrics

We quantify individual per-map accuracy, rendered perceptual accuracy, and artifacts. Per-map accuracy is computed differently depending on the semantics of the map: the roughness and specular maps are evaluated using Mean Absolute Error (L1), and the normals are evaluated using the angular distance in vector space (L<)). Further, to account for the possibility that the model always returns an accurate average value, resulting in relatively low L1, we also measure Pearson correlation ρ.

We evaluate rendered perceptual accuracy LBRDF between the ground truth stack MGT and the estimation M^ following existing metrics [Lav+21],

LBRDF=1|xy|xy1|S|(l,v)Scos2(θl)(fl,v(MGT,K)fl,v(M^,K))23                  (5.5)

where the terms are the same as in Equation 5.4.

Finally, we found that some architectural improvements in the neural network introduce artifacts in the specular and roughness maps that were not present in the input image, as illustrated in Figure 5.6. Thus, we provide an artifacts detection metric to quantify them. We start by defining a metric of homogeneity for an input image H(I),

H(I)=1|xy|xy1|d|d={,,,}FBox(I)FBox(Id)1                  (5.6)

where Id is the image shifted up, down, left, and right by a number of pixels equal to the kernel size of the box filter. Then, we define three metrics that we compute per map: e1(M)=H(M), e2(M)=H(M)H(X),ande3(M)=MI(X,M)1. X is the original input image, and MI is the Mutual Information [Rus+04], which helps discern artifacts that appear in a single map from semantic patterns (e.g., plaids, prints). Each map is labeled as having artifacts if the majority of metrics exceed their corresponding thresholds, tm(M). If any of the maps contain artifacts, the entire stack M is classified as having artifacts.

Figure 5.6: Qualitative results of some configurations of our ablation study. In red, we show that the baseline generator architecture trained on the full loss introduces artifacts, which are removed using attention on the encoder.

Figure 5.6: Qualitative results of some configurations of our ablation study. In red, we show that the baseline generator architecture trained on the full loss introduces artifacts, which are removed using attention on the encoder.

5.6 Evaluation

5.6.1 Ablation Study

In Figure 5.6 and Table 5.1, we present an ablation study to validate our model design. From a baseline U-Net [RFB15] trained with a pixel-wise loss and no data augmentation other than rescales, we add different components to the model to improve its generalization. First, we observe that using a PatchGAN [Iso+17; VPS21] discriminator provides a significant increase in accuracy. However, using a similarly-sized U-Net discriminator [SSK20], we achieve better results, particularly when using discriminator regularization. Further, Lstyle and Lfreq yield significant improvements, most notably in the normal map. To the baseline U-Net trained with the full loss, adding residual connections and one decoder per map increases accuracy. However, this setup tends to produce artifacts as it struggles to integrate global information. By adding self-attention to the encoder, we remove the artifacts; and with the MobileViT [MR21], we achieve higher quality results. Our final model contains full data augmentation, which provides small gains on generalization.

Table 5.1: Results of our ablation study, across a variety of metrics. Art. refers to our artifact detection metric. We use a color code to highlight best and worst cases.

Configuration

ρS

ρR

L<)

L1S

L1R

LBRDF

Art.↓

Baseline

0.615

0.329

7.570

0.130

0.070

0.325

0.0

Loss

Baseline +

PatchGAN

0.810

0.510

3.729

0.089

0.069

0.307

12.0

Baseline +

U-Net D.

0.854

0.658

2.950

0.085

0.068

0.299

28.0

U-Net D. +

LDdeccons

0.858

0.653

2.860

0.088

0.068

0.296

19.0

LDdeccons+

Lstyle

0.831

0.655

2.790

0.091

0.060

0.289

23.0

Lstyle+

Lfreq

0.856

0.677

2.410

0.086

0.059

0.288

20.1

Model

Lfreq+

Residual

0.855

0.665

2.310

0.089

0.060

0.285

25.5

Residual +

Decoders

0.863

0.699

2.120

0.079

0.054

0.276

18.5

Decoders +

Attention

0.860

0.665

2.080

0.080

0.059

0.275

0.5

Atention +

ViT

0.863

0.665

2.040

0.079

0.057

0.271

0.5

Augmentation

ViT +

Color

0.899

0.692

1.969

0.068

0.054

0.269

0.0

Color +

Rotations

0.870

0.682

2.050

0.078

0.055

0.271

0.2

Rotations +

Distortion

0.876

0.699

2.001

0.074

0.053

0.268

0.0

Distortion +

Erasing

0.893

0.727

1.941

0.067

0.052

0.265

0.0

5.6.2 Qualitative Analysis

We aim to understand which features are exploited by our model for making its predictions. In Figure 5.7, we show the embeddings of our generator using UMAP. It seems that our model learns to separate between material families (e.g. leathers from wovens). Interestingly, the thickness of the material is also a relevant parameter. The model is exploiting these semantic patterns without explicit supervision, providing evidence that material microgeometry plays important an important role in its optical appearance.

Figure 5.7: Embeddings of the transformer of our generator for data in the training set, reduced using UMAP [MHM18].

Figure 5.7: Embeddings of the transformer of our generator for data in the training set, reduced using UMAP [MHM18].

5.6.3 Uncertainty Evaluation

In Figure 5.8, we show the Pearson correlation matrix between the uncertainty and error for each map, our uncertainty metric (Equation 5.4), and the render perceptual metric (Equation 5.5), on our test set. As shown, neither the error nor uncertainties per map explain errors on the render space. Our proposed σBRDF achieves a remarkable correlation of 82% with the render error, validating that this metric is useful to predict errors at test time with reasonably high precision.

Figure 5.8: Top-left, correlation between render and pixel-wise losses, and uncertainties. Top-right, plot showing the correlation between our uncertainty metric and the error in render space; the renders illustrate the worst and best cases. Bottom, uncertainty and errors for different material families of the test set.

Figure 5.8: Top-left, correlation between render and pixel-wise losses, and uncertainties. Top-right, plot showing the correlation between our uncertainty metric and the error in render space; the renders illustrate the worst and best cases. Bottom, uncertainty and errors for different material families of the test set.

In the same plot at the bottom, we distill the uncertainty and render error per family, in which we observe that our model struggles to accurately and confidently predict the reflectance of satins, jacquards, and leathers more than it does for any other structure in our dataset. These structures have complex optical behavior that makes their digitization more challenging (for example, satins exhibit anisotropy, which we do not support in our material model), and are relatively uncommon in our training dataset. Figures 5.2 and 5.5 show further examples of our uncertainty estimation for a diverse set of materials.

5.6.4 Active Learning

We leverage σBRDF for active learning [Sol+21] to identify the samples that contribute most to reduce the error of the model. Figure 5.9 (top) illustrates the process. We start by training a model with 10% of the available training data and measure the uncertainties in the remainder of the dataset. We then select the samples with the highest uncertainties and retrain the model with 20%. We repeat the process with 40, 60, 80 and 100% of the training dataset. For each subset, we compare the performance of this model with four baselines: random-sampling, sampling by the highest uncertainty in normals, roughness, and specular. The results are shown in Figure 5.9 (bottom). Active sampling based on σBRDF provides significant gains in sample efficiency, obtaining better accuracy for every map. For instance, an actively trained model with our metric which uses 20% of the training data obtains comparable results to a model trained on three times more (but randomly sampled) data. While σ is typically more informative than σrough and σspec, using per-map uncertainties does not provide better results than a random strategy. Finally, Figure 5.2 shows the variation of the probabilistic samples with respect to the ground truth radiance for two materials with high and low uncertainty.

Figure 5.9: On top, illustration of our active learning algorithm. On the bottom, results of our active learning experiments. Leveraging σBRDF for actively selecting the top-k samples with the highest uncertainty, we achieve better results than a random sampling strategy for every metric we measure.

Figure 5.9: On top, illustration of our active learning algorithm. On the bottom, results of our active learning experiments. Leveraging σBRDF for actively selecting the top-k samples with the highest uncertainty, we achieve better results than a random sampling strategy for every metric we measure.

5.6.5 Comparisons with Previous Work

In Table 5.3, we compare our method with previous work on single image material capture. First, we made sure that the training data for these methods included textile materials similar to the ones we choose for testing. Emulating their capture conditions, we took images with a smartphone, with flash and ambient lighting. Note that these capture conditions are not ideal for our method, affecting the final renders if the albedo has shading gradients. However, our goal in this experiment is to evaluate the overall preservation of the material structure in the inferred maps, particularly visible in the normals.

For Shi et al. [Shi+20], we initialize the graph using a fabric material, provided in their open-source implementation and include the metallic map. Our model provides sharper and more accurate estimations without requiring optimization during test. Methods trained on style losses [Hen+21; Shi+20] degrade the semantic structure, while Zhou et al. [ZK21] generate similar albedos to ours (note that ours are captured), but provide over-smooth estimations. We provide a comparison of timings and model sizes in Table 5.2. With our efficient model design, we can provide real-time estimations without needing any optimization, which also enables our sampling-based uncertainty quantification.

Table 5.2: Model sizes for different methods, and evaluation time in seconds (mean of 100 evaluations on an RTX 2080 GPU), for different output sizes. The methods with * use test-time optimization. DiffMat [Shi+20] does not use a pre-trained model.

Method

Size (MB)

Output Dims

Eval Time (s)

Deep Inverse Rendering* [Gao+19]

167

256x256

603.5

Generative Modeling* [Hen+21]

1095.3

512x384

218.8

Diff. Material Graphs* [Shi+20]

-

512x512

1209.8

Adversarial Estimation [ZK21]

11552.2

256x256

0.078

UMat (Ours)

22.6

256x256

512x512

0.036 0.131

Table 5.3: Comparisons of our method with previous work on images captured under different lighting conditions. Top: a smartphone flash-lit image. Middle: a smartphone image with ambient light. Bottom: a flatbed scanner capture. Our method produces the best results preserving the microstructure even when capture conditions degrade due to the sensor resolution. Note that we do not estimate albedos and that absolute intensities for specular and roughness maps are not directly comparable due to differences in the material model.

IInput

Deep Inverse R. [Gao+19]

Generative Model [Hen+21]

Diff. Material Graphs [Shi+20]

Adversarial Est. [ZK21]

UMat (Ours)

5.6.6 Limitations

We show some limitations in Figure 5.10. The illumination in the scanner hides the wrinkles in the seersucker fabric, and our model predicts a flat surface. The organza fabric at the bottom is very transparent with visible holes between the yarns. Since the background is white, the model has mistakenly treated the light regions as yarn centers. For the satin at the right, the scan image exhibits specular highlights due to the directionality of the yarns. While this image may be problematic to use as an albedo, it does not affect our metrics as we use constant albedos to compute them.

Figure 5.10: Limitations cases of our method. On the left, we show a seersucker, with wrinkles that are hidden by the diffuse illumination of the device, and a translucent organza with holes between the yarns that appear very bright due to the white background of the scanner, and are therefore mistakenly treated as yarn centers. On the right, we show that for highly directional materials, such as satins, the diffuse-like illumination in our capture device sometimes introduces specular highlights.

Figure 5.10: Limitations cases of our method. On the left, we show a seersucker, with wrinkles that are hidden by the diffuse illumination of the device, and a translucent organza with holes between the yarns that appear very bright due to the white background of the scanner, and are therefore mistakenly treated as yarn centers. On the right, we show that for highly directional materials, such as satins, the diffuse-like illumination in our capture device sometimes introduces specular highlights.

5.7 Conclusions

We have presented a GAN-based method to digitize materials which leverages microgeometry appearance and a flatbed scanner as a capture device. Our method has shown better performance than state-of-the-art solutions that require a single image as input, when it comes to textile materials. To account for potential ambiguities derived from the capture setting, we have presented a method to model the uncertainty in the estimation at test time.

Managing uncertainty in machine learning projects is important to guarantee robust and functional solutions. However, this typically comes at the cost of complex or slow models. In this work, we have presented the first method to quantify uncertainty in single image material digitization, while introducing minimal impact in the training and evaluation processes. While it is currently not possible to discern the source of the uncertainty, whether it is epistemic (uncertainty which can be reduced by increasing the dataset size) or aleatoric (which is derived by a noisy data generation process), our metric has proven useful to identify ambiguous inputs, underrepresented classes, or out-of-distribution data.

We could extend our work in several ways. The most obvious extension is to estimate real albedos, so that we can deal with other types of scanning devices. Further, expanding our material model to give support for more reflectance properties, such as transmittance or anisotropy, could be useful to improve the realism in the render of textiles.

5.A Additional Implementation Details

5.A.1 Model Design

Our model is trained using a GAN framework. In this section, we detail our design choices for the generator and discriminator architecture.

Generator For the generator, we use a U-Net [RFB15] model, with a few modifications designed to maximize its efficiency, robustness, and generalization capabilities. We specify the full model architecture and layer sizes in Figure 5.11. We use residual connections [He+16; Dia+20; DLG21] in every convolutional block of the model, for better training convergence and preserving details present in the input images. We use 1 × 1 convolutions on the skip connections. Further, to maximally preserve the appearance and characteristics of every target map, we use a single decoder for each. This has been proposed in different applications, including intrinsic images, material capture, and texture synthesis [Gar+22; RG22; DLG21; ZK21]. To improve the results and enable our uncertainty metric, we append a pixel-wise MLP to each decoder in the model, with Dropout [Sri+14] regularization. Using MLPs after the decoders has been previously explored for material capture [Des+19]. We use Group Normalization [WH18] (with 16 groups per layer) and SiLU [EUD18] non-linearities throughout the model. Each convolutional block in the encoder is enhanced with a lightweight Linear Attention module [Wan+20b], with 4 attention heads, each with a dimension of 32 hidden units. On the bottleneck, we use a lightweight MobileViT Transformer block [MR21], with 128 hidden dimensions for the self-attention and MLPs, 4 layers, and a kernel size of 3. We use a Dropout [GG16] rate of 0.2. We use transposed convolutions for upsampling. Every other implementation detail in the model (strides, bias, poolings) follows [RFB15].

Figure 5.11: A full diagram of our generator, including layer sizes and output dimensions for each layer. For Self-Attention, we leverage Linear Attention [Wan+20b], we use a MobileVIT transformer on the bottleneck [MR21], Group Normalization [WH18] and SiLU [EUD18] non-linearities, one decoder per output map and residual connections in every convolutional block. In red, we show the input/output dimensions (spatial, channels) of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, regularizations, and non-linearities.

Figure 5.11: A full diagram of our generator, including layer sizes and output dimensions for each layer. For Self-Attention, we leverage Linear Attention [Wan+20b], we use a MobileVIT transformer on the bottleneck [MR21], Group Normalization [WH18] and SiLU [EUD18] non-linearities, one decoder per output map and residual connections in every convolutional block. In red, we show the input/output dimensions (spatial, channels) of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, regularizations, and non-linearities.

Discriminator For the discriminator [SSK20], we also use a U-Net [RFB15] model, with a few modifications to improve its performance as a discriminator. We specify the full model architecture and layer sizes in Figure 5.12. As in the generator, we use residual connections [He+16; Dia+20; DLG21] in every convolutional block of the model, for better training convergence and preserving details present in the input images. We use Spectral Normalization [Miy+18] and SiLU [EUD18] non-linearities throughout the model. On the bottleneck, we use a lightweight CBAM attention block [Woo+18]. The single-scalar estimation of the discriminator Denc is provided by an MLP with a similar architecture to the CBAM module. We initialize the entire model using orthogonal initialization [HXP20]. We use transposed convolutions for upsampling. Every other implementation detail in the model (strides, bias, pooling) follows [RFB15].

Figure 5.12: A full diagram of our U-Net residual discriminator, including layer sizes and output dimensions for each layer. We use a CBAM [Woo+18] module on the bottleneck and Spectral Normalization [Miy+18] throughout the network and residual connections in every convolutional block. In red, we show the input/output dimensions of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, non-linearities.

Figure 5.12: A full diagram of our U-Net residual discriminator, including layer sizes and output dimensions for each layer. We use a CBAM [Woo+18] module on the bottleneck and Spectral Normalization [Miy+18] throughout the network and residual connections in every convolutional block. In red, we show the input/output dimensions of each layer; in orange, we show attention modules; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, non-linearities.

5.A.2 Model Training

Optimization We train the models using PyTorch [Pas+19] and TorchVision [MR10]. We leverage Kornia for data augmentation [Rib+20]. To accelerate the training process, we leverage mixed precision training and automatic gradient scaling [Mic+18], and train the whole model natively on GPU. Optimization is done using Adam [KB15]. Following [SSK20], we use different learning rates for the generator lr = 0.001 and the discriminator lr = 0.005 and a batch size of 10, betas= (0.9, 0.99) and a weight decay of 0.000003. We train the models for 100 epochs, which takes 10 hours on an NVIDIA RTX 3060 GPU.

Loss Function We use the following weights for the loss function: λ<)=3,λspec=1,λrough=1,λadv=0.2,λstyle=0.25,λfreq=0.2,λcons=0.3. For the style loss function, we use the AlexNet variant of LPIPS [Zha+18], as it provides a lightweight style loss which has shown success on texture transfer [RG21].

Data Augmentation Training is done using random crops of 128 × 128 pixels. We perform random rescales uniformly on the 300, 1000 PPI range, with bilinear interpolation. We randomly rotate the SVBRDF on the [−10, 10] angle range. We use the algorithm in [RG21] to correctly rotate the normal map. On the HSV color space, we randomly change the properties of the model inputs. We randomly change the hue of the images, selected at random on the whole hue ranges. We also change the saturation, value, and contrast of the input images with factors of [0.9, 1.1], as specified on the Torchvision [MR10] ColorJitter implementation. With a probability of 0.3, we apply random Gaussian noise to the input images μ = 0, σ = 0.10. With a probability of 0.35, we apply random erasing [Zho+20] to the input images. Half of the time, we apply a random Gaussian blur to the input images σ ∼ U (0.1, 5). Finally, with a probability of 0.3, we apply cut-mix data augmentation to the discriminator, as in [SSK20].

Model Evaluation We use half precision for model evaluation, which makes the process faster and allows us to generate SVBRDF of larger sizes. For images of very high resolutions (eg ≥ 4000 × 4000), we use the stitching algorithm in [DDB20], with patch sizes of 2048 and stride of 1024 pixels.

5.A.3 Artifact Detection

The thresholds for each material map and the kernel size for the uniformity metric have been optimized given a set of 102 manually labeled textures: t1(Ms)=0.01, t2(Ms)=1.41, t3(Ms)=1.33; t1(Mr)=0.01, t2(Mr)=0.99, t3(Mr)=3.12. The size of the box filter is sBox =0.1275DPIX. The thresholds for each material map and the kernel size for the uniformity metric have been optimized given a set of 102 manually labeled textures: t1(Ms)=0.01, t2(Ms)=1.41, t3(Ms)=1.33; t1(Mr)=0.01, t2(Mr)=0.99, t3(Mr)=3.12. The size of the box filter is sBox=127.5.

5.A.4 Capture Details

We construct the training and testing dataset capturing 10 × 10 cm samples at 1000 PPI using an EPSON V850 Pro on its default settings.

For the comparisons with previous work, we place the fabric samples on a black surface. We use a Huawei Nova 5T smartphone, and capture the materials at two distances: a close-up, capturing 4 × 4 cms at a distance to the sample of 8 cms, and a full-size image, capturing 8 × 8 cms, at a distance to the sample of 13 cms. We use ISO=50, an aperture of f1.8, and a focal length of 26mm for every image. For the two distances, we capture the material using ambient illumination, with exposure times of 110 s for the full size and of 18 s for the close-up. We also capture the images using the smarphone flash lighting, with exposures of 140 s and 150 s for the full size and close-up, respectively.

5.A.5 Comparisons with Previous Work

We perform every comparison with previous work on an RTX NVIDIA 2080, using the default configuration for every method. However, for [Shi+20], we initialize the material graph with a fabric material (fabric suit vintage) provided in their repository, to better match our test data. For [Hen+21], we use the fine-tuning configuration.

5.B Dataset Analysis

Figure 5.13: Visualization of our dataset. We show the percentages of materials in our training dataset, including more detailed subcategories. On their right, we show the average specular and roughness for every category. As shown, there are some structures with distinct characteristics: Satins are highly specular due to the particularities of their yarns, and Piles (eg corduroy) or Plain Weave (eg linen fabrics) are much less glossy. We exploit this relationship between microgeometry and specularity for our estimations.

Figure 5.13: Visualization of our dataset. We show the percentages of materials in our training dataset, including more detailed subcategories. On their right, we show the average specular and roughness for every category. As shown, there are some structures with distinct characteristics: Satins are highly specular due to the particularities of their yarns, and Piles (eg corduroy) or Plain Weave (eg linen fabrics) are much less glossy. We exploit this relationship between microgeometry and specularity for our estimations.

Figure 5.14: Visualization of some Ground Truth SVBRDF of the different families in our test set. Textiles have very complex and varied microstructures which play an important role on their appearance at different scales.

Figure 5.14: Visualization of some Ground Truth SVBRDF of the different families in our test set. Textiles have very complex and varied microstructures which play an important role on their appearance at different scales.

5.C Additional Details And Results

Figure 5.15: Additional results of our uncertainty quantification method. On the top, we show a beige rib fabric with gray metallic yarns. While there is a low uncertainty for the beige yarns, the metallic yarns are harder to digitize for our model (we do not support metalness in our material model) and it shows a higher uncertainty on those yarns. In the middle, we show a leather material with a very strong structural pattern. Our model shows very low confidence for this material. In the bottom, we show a tartan fabric. Interestingly, our model shows a higher uncertainty on the blue yarns, which are less common than the other yarns in the material.

Figure 5.15: Additional results of our uncertainty quantification method. On the top, we show a beige rib fabric with gray metallic yarns. While there is a low uncertainty for the beige yarns, the metallic yarns are harder to digitize for our model (we do not support metalness in our material model) and it shows a higher uncertainty on those yarns. In the middle, we show a leather material with a very strong structural pattern. Our model shows very low confidence for this material. In the bottom, we show a tartan fabric. Interestingly, our model shows a higher uncertainty on the blue yarns, which are less common than the other yarns in the material.

Figure 5.16: Additional results of our uncertainty metric. On the top row, we show a plot between our proposed render uncertainty σBRDF and the render error, which are highly correlated. We also show average error and uncertainty per family, as well as renders with the lowest and highest errors, compared to the ground truth. Below, we show the same data for normals, specular and roughness errors and uncertainty. There are no correlations between uncertainty and errors for these maps.

Figure 5.16: Additional results of our uncertainty metric. On the top row, we show a plot between our proposed render uncertainty σBRDF and the render error, which are highly correlated. We also show average error and uncertainty per family, as well as renders with the lowest and highest errors, compared to the ground truth. Below, we show the same data for normals, specular and roughness errors and uncertainty. There are no correlations between uncertainty and errors for these maps.

Figure 5.17: Additional results of our active learning experiment. From top to bottom, we show the renders of ground truth SVBRDFs, the model trained on the 100% of the training data, a model trained on 40% of the training data available, selected following an active learning approach using our uncertainty σBRDF as guidance, and a model trained on 40% data, selected randomly. From left to right, we show a very diffuse Chiffon fabric, a highly specular Shantung fabric and a Goat Leather material with very varied microgeometry. In every case, our model trained on 40% of the data following an active learning approach obtains results which are very similar to a model trained on 100% of the data. The model trained on randomly selected 40% data produces highly inaccurate specularity and microgeometry estimations.

Figure 5.17: Additional results of our active learning experiment. From top to bottom, we show the renders of ground truth SVBRDFs, the model trained on the 100% of the training data, a model trained on 40% of the training data available, selected following an active learning approach using our uncertainty σBRDF as guidance, and a model trained on 40% data, selected randomly. From left to right, we show a very diffuse Chiffon fabric, a highly specular Shantung fabric and a Goat Leather material with very varied microgeometry. In every case, our model trained on 40% of the data following an active learning approach obtains results which are very similar to a model trained on 100% of the data. The model trained on randomly selected 40% data produces highly inaccurate specularity and microgeometry estimations.

Figure 5.18: Results of our method on datasets of previous work. On the top, we show the results of [Hen+21] on their own test set. Our method provides sharper normals which better preserve the structure of the input images. In the middle, we show the results of our method on synthetic rendered data from a material graph from [Shi+20]. Our method provides sharp normals for this synthetic image, which lies outside the distribution of our dataset, composed exclusively of real images. On the bottom, we show an albedo and normals computed using Photometric Stereo [Ike81] for a captured BTF [WGK14], for which our model also provides highly detailed results.

Figure 5.18: Results of our method on datasets of previous work. On the top, we show the results of [Hen+21] on their own test set. Our method provides sharper normals which better preserve the structure of the input images. In the middle, we show the results of our method on synthetic rendered data from a material graph from [Shi+20]. Our method provides sharp normals for this synthetic image, which lies outside the distribution of our dataset, composed exclusively of real images. On the bottom, we show an albedo and normals computed using Photometric Stereo [Ike81] for a captured BTF [WGK14], for which our model also provides highly detailed results.

Table 5.4: Comparisons of our results with previous work on images captured under different conditions for a curdoroy fabric: On the first rows, images were captured with a smartphone using the flash image. On the middle rows, using the same smartphone with ambient lighting on different scales. On the final row, a scanner image. Note that for [Shi+20] we use a fabric material for initialization and use their metallic map instead of specular, that we do not estimate albedos and that the material models are not necessarily comparable.

Input

Deep Inverse R. [Gao+19]

Generative Model [Hen+21]

Diff. Material Graphs [Shi+20]

Adversarial Est. [ZK21]

Our Method

Table 5.5: Comparisons of our results with previous work on images captured under different conditions for a suede leather: On the first rows, images were captured with a smartphone using the flash image. On the middle rows, using the same smartphone with ambient lighting on different scales. On the final row, a scanner image. Note that for [Shi+20] we use a fabric material for initialization and use their metallic map instead of specular, that we do not estimate albedos and that the material models are not necessarily comparable.

Input

Deep Inverse R. [Gao+19]

Generative Model [Hen+21]

Diff. Material Graphs [Shi+20]

Adversarial Est. [ZK21]

Our Method

BIBLIOGRAPHY

[Aba+16] Martın Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. “Tensorflow: A System for Large-Scale Machine Learning”. In: 12th Symposium on Operating Systems Design and Implementation. 2016, pp. 265–283 (cit. on p. 96).

[Aga+03] Sameer Agarwal, Ravi Ramamoorthi, Serge Belongie, and Henrik Wann Jensen. “Structured Importance Sampling of Environment Maps”. In: ACM Transactions on Graphics (TOG) 22.3 (2003), pp. 605–612 (cit. on p. 185).

[AAL16] Miika Aittala, Timo Aila, and Jaakko Lehtinen. “Reflectance Modeling by Neural Texture Synthesis”. In: ACM Transactions on Graphics (TOG) 35.4 (2016), pp. 1–13 (cit. on p. 124).

[AWL13] Miika Aittala, Tim Weyrich, and Jaakko Lehtinen. “Practical SVBRDF Capture in the Frequency Domain”. In: ACM Transactions on Graphics (TOG) 32.4 (2013), pp. 110–1 (cit. on pp. 17, 107, 128, 258).

[AWL15] Miika Aittala, Tim Weyrich, and Jaakko Lehtinen. “Two-shot SVBRDF Capture for Stationary Materials”. In: ACM Transactions on Graphics (TOG) 34.4 (2015), pp. 1–13 (cit. on pp. 17, 34, 104, 125, 258).

[Akl+18] Adib Akl, Charles Yaacoub, Marc Donias, Jean-Pierre Da Costa, and Christian Germain. “A Survey of Exemplar-Based Texture Synthesis Methods”. In: Computer Vision and Image Understanding 172 (2018), pp. 12–24 (cit. on p. 73).

[Alc+19] Raul Alcain, Carlos Heras, Iñigo Salinas, Jorge Lopez-Moreno, and Carlos Aliaga. “Microscale Optical Capture System for Digital Fabric Recreation”. In: Proceedings of the 7th International Conference on Photonics, Optics and Laser Technology - Volume 1: PHOTOPTICS, INSTICC. SciTePress, 2019, pp. 114–119 (cit. on pp. 35, 39).

[Alm+18] Amjad Almahairi, Sai Rajeshwar, Alessandro Sordoni, Philip Bachman, and Aaron Courville. “Augmented Cyclegan: Learning Many-to-Many Mappings from Unpaired Data”. In: International Conference on Machine Learning (ICML). PMLR. 2018, pp. 195–204 (cit. on p. 85).

[Ami+20] Alexander Amini, Wilko Schwarting, Ava Soleimany, and Daniela Rus. “Deep Evidential Regression”. In: Advances in Neural Information Processing Systems 33 (2020), pp. 14927–14937 (cit. on p. 126).

[AP08] Xiaobo An and Fabio Pellacini. “AppProp: All-pairs Appearance-Space Edit Propagation”. In: ACM Transactions on Graphics (TOG) 27.3 (2008), pp. 1–9 (cit. on p. 33).

[And+20] Pontus Andersson, Jim Nilsson, Tomas Akenine-Möller, Magnus Oskarsson, Kalle Åström, and Mark D Fairchild. “FLIP: A Difference Evaluator for Alternating Images.” In: Proc. ACM Comput. Graph. Interact. Tech. 3.2 (2020), pp. 15–1 (cit. on pp. 198, 199, 202).

[ASE17] Antreas Antoniou, Amos Storkey, and Harrison Edwards. “Data Augmentation Generative Adversarial Networks”. In: arXiv preprint arXiv:1711.04340 (2017) (cit. on p. 45).

[ACB19] Devansh Arpit, Vıctor Campos, and Yoshua Bengio. “How to Initialize Your Network? Robust Initialization for Weightnorm & Resnets”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on pp. 109, 119).

[ARV19] Yuki M Asano, Christian Rupprecht, and Andrea Vedaldi. “A Critical Analysis of Self-Supervision, or What We Can Learn from a Single Image”. In: International Conference on Learning Representations (ICLR). 2019 (cit. on p. 75).

[Att+22] Benjamin Attal, Jia-Bin Huang, Michael Zollhöfer, Johannes Kopf, and Changil Kim. “Learning Neural Light Fields with Ray-space Embedding”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 19819–19829 (cit. on pp. 187, 203, 212).

[Azi+19] Dejan Azinovic, Tzu-Mao Li, Anton Kaplanyan, and Matthias Nießner. “Inverse Path Tracing for Joint Material and Lighting Estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 2447–2456 (cit. on p. 203).

[Azi+23] Shekoofeh Azizi, Simon Kornblith, Chitwan Saharia, Mohammad Norouzi, and David J Fleet. “Synthetic Data from Diffusion Models Improves ImageNet Classification”. In: arXiv preprint arXiv:2304.08466 (2023) (cit. on p. 212).

[BKH16] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. “Layer Normalization”. In: arXiv preprint arXiv:1607.06450 (2016) (cit. on pp. 109, 111, 119, 176).

[Baa+22] Hendrik Baatz, Jonathan Granskog, Marios Papas, Fabrice Rousselle, and Jan Novák. “NeRF-Tex: Neural Reflectance Field Textures”. In: Computer Graphics Forum. Vol. 41. 6. Wiley Online Library. 2022, pp. 287–301 (cit. on pp. 104, 187).

[BSK23] Steve Bako, Pradeep Sen, and Anton Kaplanyan. “Deep Appearance Prefiltering”. In: ACM Transactions on Graphics (TOG) 42.2 (2023), pp. 1–23 (cit. on p. 105).

[BYK21] Wentao Bao, Qi Yu, and Yu Kong. “Evidential Deep Learning for Open Set Action Recognition”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 13349–13358 (cit. on p. 130).

[Bar+09] Connelly Barnes, Eli Shechtman, Adam Finkelstein, and Dan B Goldman. “PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing”. In: ACM Transactions on Graphics (TOG) 28.3 (2009), p. 24 (cit. on p. 74).

[Bar+15] Connelly Barnes, Fang-Lue Zhang, Liming Lou, Xian Wu, and Shi-Min Hu. “Patchtable: Efficient Patch Queries for Large Datasets and Applications”. In: ACM Transactions on Graphics (TOG) 34.4 (2015), pp. 1–10 (cit. on pp. 32, 74).

[Bar+21] Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. “Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021 (cit. on p. 186).

[Ben+21] Saguy Benaim, Ron Mokady, Amit Bermano, and Lior Wolf. “Structural Analogy from a Single Image Pair”. In: Computer Graphics Forum. Vol. 40. 1. Wiley Online Library. 2021, pp. 249–265 (cit. on pp. 31, 33, 46, 47, 56, 57, 60, 75).

[Bén+13] Pierre Bénard, Forrester Cole, Michael Kass, Igor Mordatch, James Hegarty, Martin Sebastian Senn, Kurt Fleischer, Davide Pesare, and Katherine Breeden. “Stylizing Animation by Example”. In: ACM Transactions on Graphics (TOG) 32.4 (2013), pp. 1–12 (cit. on p. 32).

[Ben+20] Gregory Benton, Marc Finzi, Pavel Izmailov, and Andrew Gordon Wilson. “Learning Invariances in Neural Networks”. In: arXiv preprint arXiv:2010.11882 (2020) (cit. on p. 43).

[BJV17] Urs Bergmann, Nikolay Jetchev, and Roland Vollgraf. “Learning Texture Manifolds with the Periodic Spatial GAN”. In: International Conference on Machine Learning (ICML). 2017, pp. 469–477 (cit. on pp. 72, 74–76, 89–91, 95, 99).

[Ber+23] Hugo Bertiche, Niloy J Mitra, Kuldeep Kulkarni, Chun-Hao Paul Huang, Tuanfeng Y Wang, Meysam Madadi, Sergio Escalera, and Duygu Ceylan. “Blowing in the Wind: CycleNet for Human Cinemagraphs from Still Images”. In: arXiv preprint arXiv:2303.08639 (2023) (cit. on pp. 17, 259).

[Bha+03] Kiran S. Bhat, Christopher D. Twigg, Jessica K. Hodgins, Pradeep K. Khosla, Zoran Popovic, and Steven M. Seitz. “Estimating Cloth Simulation Parameters from Video”. In: Symposium on Computer Animation. The Eurographics Association, 2003 (cit. on pp. 156, 158).

[Bi+18] Wenyan Bi, Peiran Jin, Hendrikje Nienborg, and Bei Xiao. “Estimating Mechanical Properties of Cloth from Videos Using Dense Motion Trajectories: Human Psychophysics and Machine Learning”. In: Journal of vision 18.5 (2018), pp. 12–12 (cit. on p. 159).

[Bie20] Lukas Biewald. Experiment Tracking with Weights and Biases. Software available from wandb.com. 2020 (cit. on pp. 108, 176, 194).

[Bit+20] Benedikt Bitterli, Chris Wyman, Matt Pharr, Peter Shirley, Aaron Lefohn, and Wojciech Jarosz. “Spatiotemporal Reservoir Resampling for Real-time Ray Tracing with Dynamic Direct Lighting”. In: ACM Transactions on Graphics (Proceedings of SIGGRAPH) 39.4 (2020) (cit. on p. 184).

[Blu+16] Adrian Blumer, Jan Novák, Ralf Habel, Derek Nowrouzezahrai, and Wojciech Jarosz. “Reduced Aggregate Scattering Operators for Path Tracing”. In: Computer Graphics Forum (Proceedings of Pacific Graphics) 35.7 (2016), pp. 461–473 (cit. on p. 203).

[Bom+21] Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. “On the Opportunities and Risks of Foundation Models”. In: (2021) (cit. on p. 159).

[Bos+20] Mark Boss, Varun Jampani, Kihwan Kim, Hendrik Lensch, and Jan Kautz. “Two-shot Spatially-Varying BRDF and Shape Estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 3982–3991 (cit. on p. 125).

[Bou+13] Katherine L Bouman, Bei Xiao, Peter Battaglia, and William T Freeman. “Estimating the Material Properties of Fabric from Video”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2013, pp. 1984–1991 (cit. on pp. 156, 159).

[BS12] Brent Burley and Walt Disney Animation Studios. “Physically-based Shading at Disney”. In: ACM SIGGRAPH. Vol. 2012. vol. 2012. 2012, pp. 1–7 (cit. on pp. 76, 127).

[Byl+22] Zoya Bylinskii, Laura Herman, Aaron Hertzmann, Stefanie Hutka, and Yile Zhang. “Towards Better User Studies in Computer Graphics and Vision”. In: arXiv preprint arXiv:2206.11461 (2022) (cit. on pp. 178, 182).

[CLA19] Carlos Castillo, Jorge López-Moreno, and Carlos Aliaga. “Recent Advances in Fabric Appearance Reproduction”. In: Computers & Graphics 84 (2019), pp. 103–121 (cit. on pp. 14, 17, 40, 256, 259).

[CHB21] Thomas Chambon, Eric Heitz, and Laurent Belcour. “Passing Multi-Channel Material Textures to a 3-Channel Loss”. In: ACM SIGGRAPH 2021 Talks. 2021, pp. 1–2 (cit. on pp. 64, 66, 67, 128).

[Cha+19] Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A. Efros. “Everybody Dance Now”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2019 (cit. on p. 85).

[Che+20a] Chengqian Che, Fujun Luan, Shuang Zhao, Kavita Bala, and Ioannis Gkioulekas. “Towards Learning-based Inverse Subsurface Scattering”. In: 2020 IEEE International Conference on Computational Photography (ICCP). IEEE. 2020, pp. 1–12 (cit. on p. 203).

[Che+22] Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. “Tensorf: Tensorial Radiance Fields”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2022, pp. 333–350 (cit. on pp. 203, 212).

[Che+17a] Dongdong Chen, Jing Liao, Lu Yuan, Nenghai Yu, and Gang Hua. “Coherent Online Video Style Transfer”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2017, pp. 1105–1114 (cit. on p. 32).

[Che+17b] Dongdong Chen, Lu Yuan, Jing Liao, Nenghai Yu, and Gang Hua. “Stylebank: An Explicit Representation for Neural Image Style Transfer”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 1897–1906 (cit. on p. 32).

[Che+20b] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. “A Simple Framework for Contrastive Learning of Visual Representations”. In: International Conference on Machine Learning (ICML). PMLR. 2020, pp. 1597–1607 (cit. on pp. 37, 45, 159).

[CXH21] Xinlei Chen, Saining Xie, and Kaiming He. “An Empirical Study of Training Self-supervised Vision Transformers”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2021, pp. 9640–9649 (cit. on p. 159).

[Che+18] Yuanqin Chen, Qian Zhang, Yaping Wu, Bo Liu, Meiyun Wang, and Yusong Lin. “Fine-tuning ResNet for Breast Cancer Classification from mammography”. In: The International Conference on Healthcare Science and Engineering. Springer. 2018, pp. 83–96 (cit. on p. 159).

[CNN22] Zhe Chen, Shohei Nobuhara, and Ko Nishino. “Invertible Neural BRDF for Object Inverse Rendering”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 44.12 (2022), pp. 9380–9395 (cit. on pp. 186, 203).

[Cho+18] Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. “Stargan: Unified Generative Adversarial Networks for Multi-domain Image-to-image Translation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 8789–8797 (cit. on p. 77).

[Cla+90] Timothy G Clapp, Hong Peng, Tushar K Ghosh, and Jeffrey W Eischen. “Indirect Measurement of the Moment-curvature Relationship for Fabrics”. In: Textile Research Journal 60.9 (1990), pp. 525–533 (cit. on pp. 156, 158).

[Cla+21] Petrik Clarberg, Wojciech Jarosz, Tomas Akenine-Möller, and Henrik Wann Jensen. “Wavelet Importance Sampling: Efficiently Evaluating Products of Complex Functions”. In: ACM Transactions on Graphics (Proceedings of SIGGRAPH) 24.3 (2021), pp. 3520–3528 (cit. on pp. 185, 189, 195, 199, 200).

[CTT17] David Clyde, Joseph Teran, and Rasmus Tamstorf. “Modeling and Data-driven Parameter Estimation for Woven Fabrics”. In: Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation. 2017, pp. 1–11 (cit. on p. 158).

[Coh+03] Michael F Cohen, Jonathan Shade, Stefan Hiller, and Oliver Deussen. “Wang Tiles for Image and Texture Generation”. In: ACM Transactions on Graphics (TOG) 22.3 (2003), pp. 287–294 (cit. on p. 72).

[CT82] Robert L Cook and Kenneth E. Torrance. “A Reflectance Model for Computer Graphics”. In: ACM Transactions on Graphics (TOG) 1.1 (1982), pp. 7–24 (cit. on p. 48).

[Dan01] Kristin J Dana. “BRDF/BTF Measurement Device”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Vol. 2. IEEE. 2001, pp. 460–466 (cit. on p. 103).

[Dan+99] Kristin J Dana, Bram Van Ginneken, Shree K Nayar, and Jan J Koenderink. “Reflectance and Texture of Real-World Surfaces”. In: ACM Transactions On Graphics (TOG) 18.1 (1999), pp. 1–34 (cit. on p. 33).

[Dar+12] Soheil Darabi, Eli Shechtman, Connelly Barnes, Dan B Goldman, and Pradeep Sen. “Image Melding: Combining Inconsistent Images Using Patch-Based Synthesis”. In: ACM Transactions on Graphics (TOG) 31.4 (2012), pp. 1–10 (cit. on p. 74).

[Dav+15] Abe Davis, Katherine L Bouman, Justin G Chen, Michael Rubinstein, Fredo Durand, and William T Freeman. “Visual Vibrometry: Estimating Material Properties from Small Motion in Video”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2015, pp. 5335–5343 (cit. on p. 159).

[De 97] Jeremy S De Bonet. “Multiresolution Sampling Procedure for Analysis and Synthesis of Texture Images”. In: Proceedings of the 24th Annual cCnference on Computer Graphics and Interactive Techniques. 1997, pp. 361–368 (cit. on p. 74).

[Dek+15] Tali Dekel, Tomer Michaeli, Michal Irani, and William T. Freeman. “Revealing and Modifying Non-Local Variations in a Single Image”. In: ACM Transactions on Graphics (TOG) (2015) (cit. on p. 93).

[DH19] Thomas Deliot and Eric Heitz. “Procedural Stochastic Textures by Tiling and Blending”. In: GPU Zen 2 (2019) (cit. on pp. 72, 74, 89–91, 95, 99).

[Den+09] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. “Imagenet: A Large-scale Hierarchical Image Database”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2009, pp. 248–255 (cit. on pp. 32, 164, 166).

[Des+18] Valentin Deschaintre, Miika Aittala, Fredo Durand, George Drettakis, and Adrien Bousseau. “Single-image SVBRDF Capture with a Rendering-Aware Deep Network”. In: ACM Transactions on Graphics (TOG) 37.4 (2018) (cit. on pp. 17, 34, 39, 54, 77, 102, 123, 125, 127, 258).

[Des+19] Valentin Deschaintre, Miika Aittala, Frédo Durand, George Drettakis, and Adrien Bousseau. “Flexible SVBRDF Capture with a Multi-Image Deep Network”. In: Computer Graphics Forum. Vol. 38. 4. Wiley Online Library. 2019, pp. 1–13 (cit. on pp. 34, 42, 72, 76, 102, 125, 141).

[DDB20] Valentin Deschaintre, George Drettakis, and Adrien Bousseau. “Guided Fine-Tuning for Large-Scale Material Transfer”. In: Computer Graphics Forum. Vol. 39. 4. Wiley Online Library. 2020, pp. 91–105 (cit. on pp. 35, 38–40, 47, 49, 50, 63, 72, 76, 105, 125, 142).

[DLG21] Valentin Deschaintre, Yiming Lin, and Abhijeet Ghosh. “Deep Polarization Imaging for 3D Shape and SVBRDF Acquisition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 15567–15576 (cit. on pp. 64, 67, 69, 127, 141).

[Dia+20] Foivos I Diakogiannis, François Waldner, Peter Caccetta, and Chen Wu. “ResUNet-a: A Deep Learning Framework for Semantic Segmentation of Remotely Sensed Data”. In: ISPRS Journal of Photogrammetry and Remote Sensing 162 (2020), pp. 94–114 (cit. on pp. 64, 66, 69, 108, 127, 141).

[Dia+15] Olga Diamanti, Connelly Barnes, Sylvain Paris, Eli Shechtman, and Olga Sorkine-Hornung. “Synthesis of Complex Image Appearance from Limited Exemplars”. In: ACM Transactions on Graphics (TOG) 34.2 (2015), pp. 1–14 (cit. on pp. 34, 104).

[Din+20] Keyan Ding, Kede Ma, Shiqi Wang, and Eero P Simoncelli. “Image Quality Assessment: Unifying Structure and Texture Similarity”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence (2020) (cit. on p. 160).

[DKB14] Laurent Dinh, David Krueger, and Yoshua Bengio. “Nice: Non-linear Independent Components Estimation”. In: arXiv preprint arXiv:1410.8516 (2014) (cit. on p. 191).

[DSB16] Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. “Density Estimation Using Real NVP”. In: arXiv preprint arXiv:1605.08803 (2016) (cit. on p. 191).

[Dod+22] Ana Dodik, Silvia Sellán, Theodore Kim, and Amanda Phillips. “Sex and Gender in the Computer Graphics Research Literature”. In: arXiv preprint arXiv:2206.00480 (2022) (cit. on pp. 178, 179, 182).

[DZP07] Weiming Dong, Ning Zhou, and Jean-Claude Paul. “Optimized Tile-Based Texture Synthesis”. In: Proceedings of Graphics Interface 2007. 2007, pp. 249–256 (cit. on p. 74).

[Don19] Yue Dong. “Deep Appearance Modeling: A Survey”. In: Visual Informatics 3.2 (2019), pp. 59–68 (cit. on pp. 34, 53).

[Don+10] Yue Dong, Jiaping Wang, Xin Tong, John Snyder, Yanxiang Lan, Moshe Ben-Ezra, and Baining Guo. “Manifold Bootstrapping for SVBRDF Capture”. In: ACM Transactions on Graphics (TOG) 29.4 (2010), pp. 1–10 (cit. on p. 34).

[DB16] Alexey Dosovitskiy and Thomas Brox. “Generating images with Perceptual Similarity Metrics based on Deep Networks”. In: Advances in Neural Information Processing Systems. 2016, pp. 658–666 (cit. on p. 74).

[DTD21] Emilien Dupont, Yee Whye Teh, and Arnaud Doucet. “Generative Models as Distributions of Functions”. In: arXiv preprint arXiv:2102.04776 (2021) (cit. on p. 75).

[Dur+19a] Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios. “Neural Spline Flows”. In: Advances in Neural Information Processing Systems. Vol. 32. 2019 (cit. on pp. 185, 191, 192, 194, 196, 197).

[Dur+19b] Conor Durkan, Artur Bekasovs, Iain Murray, and Georgios Papamakarios. “Cubic-Spline Flows”. In: Workshop on Invertible Neural Nets and Normalizing Flows: ICML 2019. 2019 (cit. on p. 191).

[Dye+18] Joanne Dyer, Diego Tamburini, Elisabeth R O’Connell, and Anna Harrison. “A Multispectral Imaging Approach Integrated into the Study of Late Antique textiles from Egypt”. In: PLoS One 13.10 (2018), e0204699 (cit. on p. 213).

[EF01] Alexei A Efros and William T Freeman. “Image Quilting for Texture Synthesis and Transfer”. In: Proceedings of the 28th annual conference on Computer Graphics and Interactive Techniques. 2001, pp. 341–346 (cit. on pp. 72, 74).

[EL99] Alexei A Efros and Thomas K Leung. “Texture Synthesis by Non-parametric Sampling”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Vol. 2. IEEE. 1999, pp. 1033–1038 (cit. on p. 74).

[EM17] Michael Elad and Peyman Milanfar. “Style Transfer via Texture Synthesis”. In: IEEE Transactions on Image Processing 26.5 (2017), pp. 2338–2351 (cit. on pp. 34, 104).

[EUD18] Stefan Elfwing, Eiji Uchibe, and Kenji Doya. “Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning”. In: Neural Networks 107 (2018), pp. 3–11 (cit. on pp. 141, 144).

[End+16] Yuki Endo, Satoshi Iizuka, Yoshihiro Kanamori, and Jun Mitani. “Deepprop: Extracting Deep Features from a Single Image for Edit Propagation”. In: Computer Graphics Forum. Vol. 35. 2. Wiley Online Library. 2016, pp. 189–201 (cit. on p. 33).

[Fan+22] Jiahui Fan, Beibei Wang, Miloš Hašan, Jian Yang, and Ling-Qi Yan. “Neural Layered BRDFs”. In: Proceedings of SIGGRAPH 2022. 2022 (cit. on pp. 116, 187).

[Fen+22] Xudong Feng, Wenchao Huang, Weiwei Xu, and Huamin Wang. “Learning-Based Bending Stiffness Parameter Estimation by a Drape Tester”. In: ACM Transactions on Graphics (TOG) 41.6 (2022), pp. 1–16 (cit. on p. 159).

[Fil+18] J. Filip, M. Kolafová, M. Havlıček, R. Vávra, M. Haindl, and Rushmeier H. “Evaluating Physical and Rendered Material Appearance”. In: The Visual Computer (Computer Graphics International 2018) (2018) (cit. on p. 109).

[FH08] Jiřı Filip and Michal Haindl. “Bidirectional Texture Function Modeling: A State of the Art Survey”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 31.11 (2008), pp. 1921–1940 (cit. on p. 103).

[FR22] Michael Fischer and Tobias Ritschel. “Plateau-Reduced Differentiable Path Tracing”. In: arXiv preprint arXiv:2211.17263 (2022) (cit. on p. 202).

[Fiš+16] Jakub Fišer, Ondřej Jamriška, Michal Lukáč, Eli Shechtman, Paul Asente, Jingwan Lu, and Daniel Sy`kora. “StyLit: Illumination-Guided Example-Based Stylization of 3D Renderings”. In: ACM Transactions on Graphics (TOG) 35.4 (2016), pp. 1–11 (cit. on p. 33).

[Fri+22] Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. “Plenoxels: Radiance Fields Without Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 5501–5510 (cit. on pp. 187, 212).

[FAW19] Anna Frühstück, Ibraheem Alhashim, and Peter Wonka. “TileGAN: Synthesis of Large-Scale Non-Homogeneous Textures”. In: ACM Transactions on Graphics (TOG) 38.4 (Apr. 2019) (cit. on pp. 34, 72, 74, 79–81, 104).

[Fu+20] Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yulan Guo, Yinghui Gao, and Biao Li. “Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs”. In: 31st British Machine Vision Conference 2020, BMVC 2020, BMVA Press, 2020 (cit. on pp. 167, 168).

[GG16] Yarin Gal and Zoubin Ghahramani. “Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning”. In: International Conference on Machine Learning (ICML). PMLR. 2016, pp. 1050–1059 (cit. on pp. 123, 126, 130, 141).

[Gal+12] Bruno Galerne, Ares Lagae, Sylvain Lefebvre, and George Drettakis. “Gabor Noise by Example”. In: ACM Transactions on Graphics (TOG) 31.4 (2012), pp. 1–9 (cit. on p. 72).

[Gao+21] Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. “Dynamic View Synthesis From Dynamic Monocular Video”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 5712–5721 (cit. on p. 186).

[Gao+19] Duan Gao, Xiao Li, Yue Dong, Pieter Peers, Kun Xu, and Xin Tong. “Deep Inverse Rendering for High-Resolution SVBRDF Estimation from an Arbitrary Number of Images”. In: ACM Transactions on Graphics (TOG) 38.4 (2019), pp. 1–15 (cit. on pp. 49, 59, 61, 125, 137, 138, 152, 153).

[GMX22] Duan Gao, Haoyuan Mu, and Kun Xu. “Neural Global Illumination: Interactive In-direct Illumination Prediction under Dynamic Area Lights”. In: IEEE Transactions on Visualization and Computer Graphics (2022) (cit. on pp. 105, 186, 203).

[Gar+14] Elena Garces, Aseem Agarwala, Diego Gutierrez, and Aaron Hertzmann. “A Similarity Measure for Illustration Style”. In: ACM Transactions on Graphics (TOG) 33.4 (2014), pp. 1–9 (cit. on pp. 157, 160, 170, 178, 180).

[Gar+23] Elena Garces, Victor Arellano, Carlos Rodriguez-Pardo, David Pascual-Hernandez, Sergio Suja, and Jorge Lopez-Moreno. “Towards Material Digitization with a Dual-scale Optical System”. In: ACM Transactions on Graphics (TOG) (2023) (cit. on pp. 15, 16, 24, 102, 109, 269).

[Gar+22] Elena Garces, Carlos Rodriguez-Pardo, Dan Casas, and Jorge Lopez-Moreno. “A Survey on Intrinsic Images: Delving Deep into Lambert and Beyond”. In: International Journal of Computer Vision (2022) (cit. on pp. 3, 4, 6, 7, 12, 22, 23, 127, 128, 141, 268).

[GES22] James Gardner, Bernhard Egger, and William Alfred Peter Smith. “Rotation-Equivariant Conditional Spherical Neural Fields for Learning a Natural Illumination Prior”. In: Advances in Neural Information Processing Systems. 2022 (cit. on pp. 19, 185, 186, 193, 196, 199, 202, 203, 260).

[Gar+19] Marc-André Gardner, Yannick Hold-Geoffroy, Kalyan Sunkavalli, Christian Gagné, and Jean-François Lalonde. “Deep Parametric Indoor Lighting Estimation”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2019, pp. 7175–7183 (cit. on p. 203).

[GEB15a] Leon Gatys, Alexander S Ecker, and Matthias Bethge. “Texture Synthesis using Convolutional Neural Networks”. In: Advances in Neural Information Processing Systems. 2015, pp. 262–270 (cit. on pp. 74, 95, 128).

[GEB15b] Leon A Gatys, Alexander S Ecker, and Matthias Bethge. “A Neural Algorithm of Artistic Style”. In: arXiv preprint arXiv:1508.06576 (2015) (cit. on pp. 32, 47, 48, 124, 128).

[GEB16a] Leon A Gatys, Alexander S Ecker, and Matthias Bethge. “Image Style Transfer using Convolutional Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2016, pp. 2414–2423 (cit. on pp. 74, 160).

[Gat+17] Leon A Gatys, Alexander S Ecker, Matthias Bethge, Aaron Hertzmann, and Eli Shechtman. “Controlling Perceptual Factors in Neural Style Transfer”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 3985–3993 (cit. on pp. 32, 74).

[GEB16b] Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. “Image Style Transfer Using Convolutional Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vol. 2016-Decem. June 2016, pp. 2414–2423 (cit. on p. 79).

[Gau+22] Alban Gauthier, Robin Faury, Jérémy Levallois, Theo Thonat, Jean-Marc Thiery, and Tamy Boubekeur. “MIPNet: Neural Normal-to-Anisotropic-Roughness MIP mapping”. In: ACM Transactions on Graphics (TOG) 41.6 (2022), pp. 1–12 (cit. on p. 105).

[Gaw+21] Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, et al. “A Survey of Uncertainty in Deep Neural Networks”. In: arXiv preprint arXiv:2107.03342 (2021) (cit. on p. 126).

[Geo+18] Iliyan Georgiev, Thiago Ize, Mike Farnsworth, Ramón Montoya-Vozmediano, Alan King, Brecht Van Lommel, Angel Jimenez, Oscar Anson, Shinji Ogaki, Eric Johnston, et al. “Arnold: A Brute-force Production Path Tracer”. In: ACM Transactions on Graphics (TOG) 37.3 (2018), pp. 1–12 (cit. on p. 48).

[Gil+14] Guillaume Gilet, Basile Sauvage, Kenneth Vanhoey, Jean-Michel Dischler, and Djamchid Ghazanfarpour. “Local Random-phase Noise for Procedural Texturing”. In: ACM Transactions on Graphics (TOG) 33.6 (2014), pp. 1–11 (cit. on p. 72).

[Goo+14a] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. “Generative Adversarial Nets”. In: Advances in Neural Information Processing Systems. 2014, pp. 2672–2680 (cit. on p. 39).

[Goo+14b] Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. “Generative Adversarial Nets”. In: Advances in Neural Information Processing Systems. Vol. 3. January. Neural information processing systems foundation, June 2014, pp. 2672–2680 (cit. on p. 79).

[Gri+03] Eitan Grinspun, Anil N Hirani, Mathieu Desbrun, and Peter Schröder. “Discrete Shells”. In: Proceedings of the 2003 ACM SIGGRAPH/Eurographics symposium on Computer animation. Citeseer. 2003, pp. 62–67 (cit. on pp. 161, 162).

[Gro+20] Aditya Grover, Christopher Chute, Rui Shu, Zhangjie Cao, and Stefano Ermon. “Align-flow: Cycle Consistent Learning from Multiple Domains via Normalizing Flows”. In: Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 34. 04. 2020, pp. 4028–4035 (cit. on p. 186).

[Gu+18] Shuyang Gu, Congliang Chen, Jing Liao, and Lu Yuan. “Arbitrary Style Transfer with Deep Feature Reshuffle”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 8222–8231 (cit. on p. 32).

[Gua+16] Darya Guarnera, Giuseppe Claudio Guarnera, Abhijeet Ghosh, Cornelia Denk, and Mashhuda Glencross. “BRDF Representation and Acquisition”. In: Computer Graphics Forum. Vol. 35. 2. Wiley Online Library. 2016, pp. 625–650 (cit. on p. 34).

[Gua+19] Giuseppe Claudio Guarnera, Dar’ya Guarnera, Gregory J Ward, Mashhuda Glencross, and Ian Hall. “BxDF Material Acquisition, Representation, and Rendering for VR and Design”. In: SIGGRAPH Asia 2019 Courses. 2019, pp. 1–21 (cit. on p. 3).

[Gue+20] P Guehl, R Allègre, J-M Dischler, B Benes, and E Galin. “Semi-Procedural Textures Using Point Process Texture Basis Functions”. In: Computer Graphics Forum. Vol. 39. 4. Wiley Online Library. 2020, pp. 159–171 (cit. on pp. 34, 72, 92, 104).

[Gui+17] Geoffrey Guingo, Basile Sauvage, Jean-Michel Dischler, and Marie-Paule Cani. “Bilayer Textures: A Model for Synthesis and Deformation of Composite Textures”. In: Computer Graphics Forum. Vol. 36. 4. 2017, pp. 111–122 (cit. on p. 72).

[Guo+21] Jie Guo, Shuichang Lai, Chengzhi Tao, Yuelong Cai, Lei Wang, Yanwen Guo, and Ling-Qi Yan. “Highlight-Aware Two-Stream Network for Single-Image SVBRDF Acquisition”. In: ACM Transactions on Graphics (TOG) 40.4 (2021), pp. 1–14 (cit. on pp. 123, 125).

[Guo+20a] Yu Guo, Miloš Hašan, Lingqi Yan, and Shuang Zhao. “A Bayesian Inference Framework for Procedural Material Parameter Estimation”. In: Computer Graphics Forum. Vol. 39. 7. Wiley Online Library. 2020, pp. 255–266 (cit. on pp. 125, 126).

[Guo+20b] Yu Guo, Cameron Smith, Miloš Hašan, Kalyan Sunkavalli, and Shuang Zhao. “MaterialGAN: Reflectance Capture using a Generative SVBRDF Model”. In: ACM Transactions on Graphics (TOG) 39.6 (2020), 254:1–254:13 (cit. on pp. 17, 35, 42, 72, 76, 102, 123, 125, 258).

[Guo+19] Yunhui Guo, Honghui Shi, Abhishek Kumar, Kristen Grauman, Tajana Rosing, and Rogerio Feris. “Spottune: Transfer Learning through Adaptive Fine-tuning”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4805–4814 (cit. on p. 159).

[HDL16] David Ha, Andrew Dai, and Quoc V Le. “Hypernetworks”. In: arXiv preprint arXiv:1609.09106 (2016) (cit. on p. 75).

[HMV09] Simon Haegler, Pascal Müller, and Luc Van Gool. “Procedural Modeling for Digital Cultural Heritage”. In: EURASIP Journal on Image and Video Processing 2009 (2009), pp. 1–1 (cit. on p. 213).

[Hai12] Vávra R. Haindl M. Filip J. “Digital Material Appearance: the Curse of Tera-Bytes”. In: ERCIM News 90 (2012), pp. 49–50 (cit. on pp. 108, 109, 111).

[Ham+21] Hendrik Hameeuw, Godelieve Watteeuw, Bruno Vandermeulen, and Marc Proesmans. “The Painted Panels of the Early Sixteenth Century Mechelen Enclosed Gardens Art Technical Examination with the Photometric Stereo: White Light and Multispectral Microdomes”. In: Papers Presented at the Twentieth Symposium for the Study of Underdrawing and Technology in Painting held in Mechelen and Leuven, 11-13 January 2017. Vol. 20. Peeters Publishers; Leuven. 2021, pp. 165–175 (cit. on p. 213).

[Han+22a] Jiyeon Han, Hwanil Choi, Yunjey Choi, Junho Kim, Jung-Woo Ha, and Jaesik Choi. “Rarity Score: A New Metric to Evaluate the Uncommonness of Synthesized Images”. In: arXiv preprint arXiv:2206.08549 (2022) (cit. on p. 128).

[Han+22b] Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. “A Survey on Vision Transformer”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence (2022) (cit. on p. 164).

[Haq+23] Ayaan Haque, Matthew Tancik, Alexei Efros, Aleksander Holynski, and Angjoo Kanazawa. “Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions”. In: arXiv preprint arXiv:2303.12789 (2023) (cit. on p. 116).

[HHM22] Jon Hasselgren, Nikolai Hofmann, and Jacob Munkberg. “Shape, Light & Material Decomposition from Images Using Monte Carlo Rendering and Denoising”. In: arXiv preprint arXiv:2206.03380 (2022) (cit. on p. 203).

[HFM10] Vlastimil Havran, Jirı Filip, and Karol Myszkowski. “Bidirectional Texture Function Compression Based on Multi-level Vector Quantization”. In: Computer Graphics Forum. Vol. 29. 1. Wiley Online Library. 2010, pp. 175–190 (cit. on p. 103).

[He+22] Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. “Masked Autoencoders are Scalable Vision Learners”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 16000–16009 (cit. on p. 159).

[He+16] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. “Deep Residual Learning for Image Recognition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vol. 2016-Decem. 2016, pp. 770–778 (cit. on pp. 64, 66, 69, 78, 94, 108, 109, 127, 141, 164, 176).

[He+18] Mingming He, Dongdong Chen, Jing Liao, Pedro V Sander, and Lu Yuan. “Deep Exemplar-Based Colorization”. In: ACM Transactions on Graphics (TOG) 37.4 (2018), pp. 1–16 (cit. on p. 33).

[He+19] Mingming He, Jing Liao, Dongdong Chen, Lu Yuan, and Pedro V Sander. “Progressive Color Transfer with Dense Semantic Correspondences”. In: ACM Transactions on Graphics (TOG) 38.2 (2019), pp. 1–18 (cit. on p. 33).

[HB95] David J Heeger and James R Bergen. “Pyramid-based Texture Analysis/Synthesis”. In: Proceedings of the 22nd Annual Conference on Computer Graphics and Interactive Techniques. 1995, pp. 229–238 (cit. on p. 74).

[Hei20] Eric Heitz. “Can’t Invert the CDF? The Triangle-Cut Parameterization of the Region under the Curve”. In: Computer Graphics Forum 39.4 (2020), pp. 121–132 (cit. on p. 185).

[HN18] Eric Heitz and Fabrice Neyret. “High-Performance By-Example Noise Using a Histogram-Preserving Blending Operator”. In: Proceedings of the ACM on Computer Graphics and Interactive Techniques 1.2 (2018), pp. 1–25 (cit. on p. 72).

[Hei+21] Eric Heitz, Kenneth Vanhoey, Thomas Chambon, and Laurent Belcour. “A Sliced Wasserstein Loss for Neural Texture Synthesis”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 9412–9420 (cit. on pp. 75, 93).

[Hel+20] Leonhard Helminger, Abdelaziz Djelouah, Markus Gross, and Christopher Schroers. “Lossy Image Compression with Normalizing Flows”. In: arXiv preprint arXiv:2008.10486 (2020) (cit. on p. 186).

[HG16] Dan Hendrycks and Kevin Gimpel. “Gaussian Error Linear Units (GELUs)”. In: arXiv preprint arXiv:1606.08415 (2016) (cit. on pp. 109, 119).

[Hen+21] Philipp Henzler, Valentin Deschaintre, Niloy J Mitra, and Tobias Ritschel. “Generative Modelling of BRDF Textures from Flash Images”. In: ACM Transactions on Graphics (Proc. SIGGRAPH Asia) 40.6 (2021) (cit. on pp. 17, 76, 102, 123, 124, 137, 138, 143, 151–153, 258).

[Her+20] Amir Hertz, Rana Hanocka, Raja Giryes, and Daniel Cohen-Or. “Deep Geometric Texture Synthesis”. In: ACM Transactions on Graphics (TOG) 39.4 (2020), pp. 108–1 (cit. on p. 72).

[Her+01] Aaron Hertzmann, Charles E Jacobs, Nuria Oliver, Brian Curless, and David H Salesin. “Image Analogies”. In: Proceedings of the 28th annual conference on Computer graphics and interactive techniques. 2001, pp. 327–340 (cit. on pp. 31, 32, 46, 56, 57, 60).

[HS05] Aaron Hertzmann and Steven M Seitz. “Example-based Photometric Stereo: Shape Reconstruction with General, Varying BRDFs”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 27.8 (2005), pp. 1254–1264 (cit. on p. 34).

[HS03] Aaron Hertzmann and Steven M Seitz. “Shape and Materials by example: A Photometric Stereo Approach”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vol. 1. IEEE. 2003, pp. I–I (cit. on p. 34).

[Heu+17] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. “Gans Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium”. In: arXiv preprint arXiv:1706.08500 (2017) (cit. on p. 85).

[Hil+15] Stephen Hill, Stephen McAuley, Brent Burley, et al. “Physically Based Shading in Theory and Practice”. In: ACM SIGGRAPH 2015 Courses. SIGGRAPH ’15. Los Angeles, California: Association for Computing Machinery, 2015 (cit. on p. 11).

[HS06] Geoffrey E Hinton and Ruslan R Salakhutdinov. “Reducing the Dimensionality of Data with Neural Networks”. In: science 313.5786 (2006), pp. 504–507 (cit. on p. 103).

[Hin+21] Tobias Hinz, Matthew Fisher, Oliver Wang, and Stefan Wermter. “Improved Techniques for Training Single-Image GANs”. In: Proceedings - 2021 IEEE Winter Conference on Applications of Computer Vision, WACV 2021. Jan. 2021, pp. 1300–1309 (cit. on p. 75).

[HJA20] Jonathan Ho, Ajay Jain, and Pieter Abbeel. “Denoising Diffusion Probabilistic Models”. In: Advances in Neural Information Processing Systems 33 (2020), pp. 6840–6851 (cit. on pp. 127, 209).

[HS13] Shahera Hossain and Seiichi Serikawa. “Texture Databases–A Comprehensive Survey”. In: Pattern Recognition Letters 34.15 (2013), pp. 2007–2022 (cit. on p. 109).

[Hu+20] Bingyang Hu, Jie Guo, Yanjun Chen, Mengtian Li, and Yanwen Guo. “DeepBRDF: A Deep Representation for Manipulating Measured BRDF”. In: Computer Graphics Forum. Vol. 39. 2. Wiley Online Library. 2020, pp. 157–166 (cit. on pp. 105, 203).

[Hu+22a] Dongting Hu, Liuhua Peng, Tingjin Chu, Xiaoxing Zhang, Yinian Mao, Howard Bondell, and Mingming Gong. “Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2022 (cit. on pp. 126, 130).

[HSS18] Jie Hu, Li Shen, and Gang Sun. “Squeeze-and-excitation Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 7132–7141 (cit. on p. 164).

[HXP20] Wei Hu, Lechao Xiao, and Jeffrey Pennington. “Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks”. In: arXiv preprint arXiv:2001.05992 (2020) (cit. on p. 141).

[HDR19] Yiwei Hu, Julie Dorsey, and Holly Rushmeier. “A Novel Framework for Inverse Procedural Texture Modeling”. In: ACM Transactions on Graphics (TOG) 38.6 (2019), pp. 1–14 (cit. on p. 72).

[Hu+22b] Yiwei Hu, Miloš Hašan, Paul Guerrero, Holly Rushmeier, and Valentin Deschaintre. “Controlling Material Appearance by Examples”. In: Computer Graphics Forum. Vol. 41. 4. Wiley Online Library. 2022, pp. 117–128 (cit. on p. 105).

[Hu+22c] Yiwei Hu, Chengan He, Valentin Deschaintre, Julie Dorsey, and Holly Rushmeier. “An Inverse Procedural Modeling Pipeline for SVBRDF Maps”. In: ACM Transactions on Graphics (TOG) 41.2 (2022), pp. 1–17 (cit. on p. 125).

[Hu+19] Yuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun, Nathan Carr, Jonathan Ragan-Kelley, and Fredo Durand. “Diff Taichi: Differentiable Programming for Physical Simulation”. In: International Conference on Learning Representations (ICLR). 2019 (cit. on p. 158).

[Hua+18a] Chin-Wei Huang, David Krueger, Alexandre Lacoste, and Aaron Courville. “Neural Autoregressive Flows”. In: International Conference on Machine Learning (ICML). PMLR. 2018, pp. 2078–2087 (cit. on p. 191).

[HB17] Xun Huang and Serge Belongie. “Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2017, pp. 1501–1510 (cit. on p. 32).

[Hua+18b] Xun Huang, Ming-Yu Liu, Serge Belongie, and Jan Kautz. “Multimodal Unsupervised Image-to-image Translation”. In: Proceedings of the European Conference on Computer Vision (ECCV). Sept. 2018 (cit. on p. 85).

[Hua+19] Zhiyuan Huang, Mansur Arief, Henry Lam, and Ding Zhao. “Evaluation Uncertainty in Data-Driven Self-Driving Testing”. In: 2019 IEEE Intelligent Transportation Systems Conference (ITSC). IEEE. 2019, pp. 1902–1907 (cit. on p. 126).

[HEW17] Markus Huber, Bernhard Eberhardt, and Daniel Weiskopf. “Cloth Animation Retrieval Using a Motion-shape Signature”. In: IEEE Computer Graphics and Applications 37.6 (2017), pp. 52–64 (cit. on p. 159).

[HZ21] Drew A Hudson and C Lawrence Zitnick. “Generative Adversarial Transformers”. In: arXiv preprint arXiv:2103.01209 (2021) (cit. on p. 93).

[Hwa+22] Inseung Hwang, Daniel S Jeon, Adolfo Muñoz, Diego Gutierrez, Xin Tong, and Min H Kim. “Sparse Ellipsometry: Portable Acquisition of Polarimetric SVBRDF and Shape with Unstructured Flash Photography”. In: ACM Transactions on Graphics (TOG) 41.4 (2022), pp. 1–14 (cit. on p. 129).

[Ike81] Katsushi Ikeuchi. “Determining Surface Orientations of Specular Surfaces by Using the Photometric Stereo Method”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 6 (1981), pp. 661–669 (cit. on pp. 34, 40, 50, 84, 112, 151).

[IS15] Sergey Ioffe and Christian Szegedy. “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”. In: International Conference on Machine Learning (ICML). PMLR. 2015, pp. 448–456 (cit. on pp. 54, 64, 194).

[Iso+17] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. “Image-to-Image Translation with Conditional Adversarial Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 5967–5976 (cit. on pp. 38, 39, 54, 56, 78, 79, 94, 107, 128, 132, 193).

[Jak+22] Wenzel Jakob, Sébastien Speierer, Nicolas Roussel, Merlin Nimier-David, Delio Vicini, Tizian Zeltner, Baptiste Nicolet, Miguel Crespo, Vincent Leroy, and Ziyi Zhang. Mitsuba 3 Renderer. Version 3.1.1. https://mitsuba-renderer.org. 2022 (cit. on pp. 189, 202).

[Jam+19] Ondřej Jamriška, Šárka Sochorová, Ondřej Texler, Michal Lukáč, Jakub Fišer, Jingwan Lu, Eli Shechtman, and Daniel Sy`kora. “Stylizing Video by Example”. In: ACM Transactions on Graphics (TOG) 38.4 (2019), pp. 1–11 (cit. on p. 32).

[Jan+17] Michael Janner, Jiajun Wu, Tejas D Kulkarni, Ilker Yildirim, and Josh Tenenbaum. “Self-Supervised Intrinsic Image Decomposition”. In: Advances in Neural Information Processing Systems. 2017, pp. 5936–5946 (cit. on pp. 64, 67, 69, 78).

[JBH19] Miguel Jaques, Michael Burke, and Timothy Hospedales. “Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video”. In: International Conference on Learning Representations (ICLR). 2019 (cit. on p. 158).

[JCJ09] Wojciech Jarosz, Nathan A. Carr, and Henrik Wann Jensen. “Importance Sampling Spherical Harmonics”. In: Computer Graphics Forum (Proceedings of Eurographics) 28.2 (2009), pp. 577–586 (cit. on p. 185).

[JBV16] Nikolay Jetchev, Urs Bergmann, and Roland Vollgraf. “Texture Synthesis with Spatial Generative Adversarial Networks”. In: arXiv preprint arXiv:1611.08207 (2016) (cit. on pp. 72, 74).

[Jia+21] Liming Jiang, Bo Dai, Wayne Wu, and Chen Change Loy. “Focal Frequency Loss for Image Reconstruction and Synthesis”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 13919–13929 (cit. on pp. 107, 128).

[Jin+22] Wenhua Jin, Beibei Wang, Miloš Hašan, Yu Guo, Steve Marschner, and Ling-Qi Yan. “Woven Fabric Capture from a Single Photo”. In: Proceedings of SIGGRAPH Asia 2022. 2022 (cit. on p. 125).

[Jin+19] Yongcheng Jing, Yezhou Yang, Zunlei Feng, Jingwen Ye, Yizhou Yu, and Mingli Song. “Neural Style Transfer: A review”. In: IEEE Transactions on Visualization and Computer Graphics (2019) (cit. on p. 32).

[JAF16] Justin Johnson, Alexandre Alahi, and Li Fei-Fei. “Perceptual Losses for Real-time Style Transfer and Super-resolution”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2016, pp. 694–711 (cit. on pp. 32, 74).

[Jos+22] Laurent Valentin Jospin, Hamid Laga, Farid Boussaid, Wray Buntine, and Mohammed Bennamoun. “Hands-on Bayesian Neural Networks—A Tutorial for Deep Learning Users”. In: IEEE Computational Intelligence Magazine 17.2 (2022), pp. 29–48 (cit. on p. 126).

[JC20] Eunjung Ju and Myung Geol Choi. “Estimating Cloth Simulation Parameters From a Static Drape Using Neural Networks”. In: IEEE Access 8 (2020), pp. 195113–195121 (cit. on pp. 156, 159).

[KBH03] Florian Kainz, Rod Bogart, and Drew Hess. “The OpenEXR Image File Format”. In: ACM SIGGRAPH Technical Sketches (2003) (cit. on p. 185).

[Kaj86] James T Kajiya. “The Rendering Equation”. In: Proceedings of the 13th annual conference on Computer graphics and interactive techniques. 1986, pp. 143–150 (cit. on pp. 4, 8).

[Kam+16] Christos Kampouris, Stefanos Zafeiriou, Abhijeet Ghosh, and Sotiris Malassiotis. “Fine-Grained Material Classification Using Micro-Geometry and Reflectance”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2016, pp. 778–792 (cit. on p. 40).

[KG13] Brian Karis and Epic Games. “Real Shading in Unreal Engine 4”. In: Proc. Physically Based Shading Theory Practice 4.3 (2013), p. 1 (cit. on p. 127).

[Kar+18] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. “Progressive Growing of GANs for Improved Quality, Stability, and Variation”. In: International Conference on Learning Representations (ICLR). 2018 (cit. on pp. 74, 80).

[Kar+20a] Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. “Training Generative Adversarial Networks with Limited Data”. In: arXiv preprint arXiv:2006.06676 (2020) (cit. on pp. 36, 109).

[Kar+21] Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. “Alias-free Generative Adversarial Networks”. In: Advances in Neural Information Processing Systems. Vol. 34. 2021, pp. 852–863 (cit. on p. 186).

[KLA19] Tero Karras, Samuli Laine, and Timo Aila. “A Style-Based Generator Architecture for Generative Adversarial Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4401–4410 (cit. on p. 74).

[Kar+20b] Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. “Analyzing and Improving the Image Quality of StyleGAN”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2020 (cit. on pp. 74, 85).

[Kas+15] Alexandre Kaspar, Boris Neubert, Dani Lischinski, Mark Pauly, and Johannes Kopf. “Self Tuning Texture Optimization”. In: Computer Graphics Forum. Vol. 34. 2. 2015, pp. 349–359 (cit. on p. 74).

[KZP19] Sergey Kastryulin, Dzhamil Zakirov, and Denis Prokopenko. PyTorch Image Quality: Metrics and Measure for Image Quality Assessment. Open-source software available at https://github.com/photosynthesis-team/piq. 2019 (cit. on pp. 182, 198).

[Kau17] Eric Kauderer-Abrams. “Quantifying Translation-invariance in Convolutional Neural Networks”. In: arXiv preprint arXiv:1801.01450 (2017) (cit. on p. 37).

[Kau+00] Jan Kautz, Pere-Pau Vázquez, Wolfgang Heidrich, and Hans-Peter Seidel. “Unified Approach to Prefiltered Environment Maps”. In: Proceedings of the Eurographics Workshop on Rendering Techniques 2000. 2000, pp. 185–196 (cit. on p. 185).

[Kaw80] Sueo Kawabata. “The Standardization and Analysis of Hand Evaluation”. In: The Textile Machinery Society of Japan (1980) (cit. on pp. 156, 158).

[KG17] Alex Kendall and Yarin Gal. “What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?” In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 126).

[Khe+21] Ilyes Khemakhem, Ricardo Monti, Robert Leech, and Aapo Hyvarinen. “Causal Autoregressive Flows”. In: International Conference on Artificial Intelligence and Statistics. PMLR. 2021, pp. 3520–3528 (cit. on p. 191).

[KB15] Diederik P. Kingma and Jimmy Ba. “Adam: A Method for Stochastic Optimization”. In: International Conference on Learning Representations (ICLR). 2015 (cit. on pp. 39, 55, 67, 83, 94, 110, 142, 194).

[Kin+16] Durk P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. “Improved Variational Inference with Inverse Autoregressive Flow”. In: Advances in Neural Information Processing Systems 29 (2016) (cit. on p. 191).

[KD18] Durk P. Kingma and Prafulla Dhariwal. “Glow: Generative Flow with Invertible 1x1 Convolutions”. In: Advances in Neural Information Processing Systems. Vol. 31. 2018 (cit. on pp. 186, 191).

[Kir+23] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. “Segment Anything”. In: arXiv preprint arXiv:2304.02643 (2023) (cit. on p. 212).

[KPB20] Ivan Kobyzev, Simon JD Prince, and Marcus A Brubaker. “Normalizing Flows: An Introduction and Review of Current Methods”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 43.11 (2020), pp. 3964–3979 (cit. on p. 191).

[Kol+20] Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. “Big Transfer (bit): General Visual Representation Learning”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2020, pp. 491–507 (cit. on p. 159).

[KSL19] Simon Kornblith, Jonathon Shlens, and Quoc V Le. “Do Better Imagenet Models Transfer Better?” In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 2661–2671 (cit. on p. 159).

[Kou+03] Melissa L Koudelka, Sebastian Magda, Peter N Belhumeur, and David J Kriegman. “Acquisition, Compression, and Synthesis of Bidirectional Texture Functions”. In: 3rd International Workshop on Texture Analysis and Synthesis (Texture 2003). 2003, pp. 59–64 (cit. on p. 103).

[Kry+21] Michael C Krygier, Tyler LaBonte, Carianne Martinez, Chance Norris, Krish Sharma, Lincoln N Collins, Partha P Mukherjee, and Scott A Roberts. “Quantifying the Unknown Impact of Segmentation Uncertainty on Image-Based Simulations”. In: Nature Communications 12.1 (2021), pp. 1–11 (cit. on p. 126).

[KLG20] Sandra Kuijpers, Christiane Luible-Bär, and R. Hugh Gong. “The Measurement of Fabric Properties for Virtual Simulation—A Critical Review”. In: IEEE SA INDUSTRY CONNECTIONS (2020), pp. 1–43 (cit. on p. 158).

[KL51] Solomon Kullback and Richard A Leibler. “On Information and Sufficiency”. In: The Annals of Mathematical Statistics 22.1 (1951), pp. 79–86 (cit. on pp. 196, 197).

[Kum+22] Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. “Fine-tuning Can Distort Pretrained Features and Underperform Out-of-distribution”. In: arXiv preprint arXiv:2202.10054 (2022) (cit. on p. 159).

[Kum+19] Manoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn, Sergey Levine, Laurent Dinh, and Durk Kingma. “Videoflow: A Flow-based Generative Model for Video”. In: arXiv preprint arXiv:1903.01434 2.5 (2019), p. 3 (cit. on p. 186).

[Kur+19] Karol Kurach, Mario Lučić, Xiaohua Zhai, Marcin Michalski, and Sylvain Gelly. “A Large-Scale Study on Regularization and Normalization in GANs”. In: International Conference on Machine Learning (ICML). PMLR. 2019, pp. 3581–3590 (cit. on pp. 77, 94).

[Kur+22] Alexander Kurz, Katja Hauser, Hendrik Alexander Mehrtens, Eva Krieghoff-Henning, Achim Hekler, Jakob Nikolas Kather, Stefan Fröhling, Christof von Kalle, Titus Josef Brinker, et al. “Uncertainty Estimation in Medical Image Classification: Systematic Review”. In: JMIR Medical Informatics 10.8 (2022) (cit. on p. 126).

[Kuz+21] Alexandr Kuznetsov, Krishna Mullia, Zexiang Xu, Miloš Hašan, and Ravi Ramamoorthi. “NeuMIP: Multi-Resolution Neural Materials”. In: ACM Transactions on Graphics (TOG) 40.4 (2021), pp. 1–13 (cit. on pp. 18, 102, 103, 106, 107, 109–116, 119, 120, 187, 207, 259, 264).

[Kuz+22] Alexandr Kuznetsov, Xuezheng Wang, Krishna Mullia, Fujun Luan, Zexiang Xu, Milos Hasan, and Ravi Ramamoorthi. “Rendering Neural Materials on Curved Surfaces”. In: ACM SIGGRAPH 2022 Conference Proceedings. 2022, pp. 1–9 (cit. on pp. 102, 104, 116, 187).

[Kwa+05] Vivek Kwatra, Irfan Essa, Aaron Bobick, and Nipun Kwatra. “Texture Optimization for Example-Based Synthesis”. In: ACM Transactions on Graphics (TOG). Vol. 24. 3. ACM. 2005, pp. 795–802 (cit. on pp. 72, 74).

[Kwa+03] Vivek Kwatra, Arno Schödl, Irfan Essa, Greg Turk, and Aaron Bobick. “Graphcut Textures: Image and Video Synthesis Ysing Graph Cuts”. In: ACM Transactions on Graphics (TOG) 22.3 (2003), pp. 277–286 (cit. on p. 74).

[LGG19] Manuel Lagunas, Elena Garces, and Diego Gutierrez. “Learning Icons Appearance Similarity”. In: Multimedia Tools and Applications 78.8 (2019), pp. 10733–10751 (cit. on p. 160).

[Lag+19] Manuel Lagunas, Sandra Malpica, Ana Serrano, Elena Garces, Diego Gutierrez, and Belen Masia. “A Similarity Measure for Material Appearance”. In: ACM Transactions on Graphics (TOG) 38.4 (2019), pp. 1–12 (cit. on pp. 157, 160, 180).

[LKA13] Samuli Laine, Tero Karras, and Timo Aila. “Megakernels Considered Harmful: Wavefront Path Tracing on GPUs”. In: Proceedings of the 5th High-Performance Graphics Conference. 2013, pp. 137–143 (cit. on pp. 109, 195).

[LPB17] Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. “Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on pp. 126, 130).

[LS98] Greg Ward Larson and Rob Shakespeare. Rendering with Radiance: The Art and Science of Lighting Visualization. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1998 (cit. on p. 185).

[Lav+21] Guillaume Lavoué, Nicolas Bonneel, Jean-Philippe Farrugia, and Cyril Soler. “Perceptual Quality of BRDF Approximations: Dataset and Metrics”. In: Computer Graphics Forum. Vol. 40. 2. Wiley Online Library. 2021, pp. 327–338 (cit. on pp. 130, 131, 193).

[Led+17] Christian Ledig, Lucas Theis, Ferenc Huszár, et al. “Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2017-Janua (Sept. 2017), pp. 105–114 (cit. on pp. 79, 94).

[LH06] Sylvain Lefebvre and Hugues Hoppe. “Appearance-Space Texture Synthesis”. In: ACM Transactions on Graphics (TOG) 25.3 (2006), pp. 541–548 (cit. on pp. 34, 104).

[Lei+21] Zhao Lei, Yi Zeng, Peng Liu, and Xiaohui Su. “Active Deep Learning for Hyperspectral Image Classification with Uncertainty Learning”. In: IEEE Geoscience and Remote Sensing Letters 19 (2021) (cit. on p. 126).

[LM01] Thomas Leung and Jitendra Malik. “Representing and Recognizing the Visual Appearance of Materials Using Three-Dimensional Textons”. In: International Journal of Computer Vision 43.1 (2001), pp. 29–44 (cit. on pp. 33, 39).

[LLB22] Wei-Hong Li, Xialei Liu, and Hakan Bilen. “Cross-domain Few-shot Learning with Task-specific Adapters”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, pp. 7161–7170 (cit. on p. 159).

[Li+17a] Xiao Li, Yue Dong, Pieter Peers, and Xin Tong. “Modeling Surface Appearance from a Single Photograph Using Self-Augmented Convolutional Neural Networks”. In: ACM Transactions on Graphics (TOG) 36.4 (2017), pp. 1–11 (cit. on pp. 34, 124).

[Li+19] Xiaoyu Li, Bo Zhang, Pedro V Sander, and Jing Liao. “Blind Geometric Distortion Correction on Images through Deep Learning”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4855–4864 (cit. on p. 93).

[Li+22] Yifei Li, Tao Du, Kui Wu, Jie Xu, and Wojciech Matusik. “DiffCloth: Differentiable Cloth Simulation with Dry Frictional Contact”. In: ACM Transactions on Graphics (TOG) (2022) (cit. on p. 158).

[Li+17b] Yijun Li, Chen Fang, Jimei Yang, Zhaowen Wang, Xin Lu, and Ming-Hsuan Yang. “Universal Style Transfer via Feature Transforms”. In: Advances in Neural Information Processing Systems. 2017, pp. 386–396 (cit. on p. 32).

[LAA08] Yuanzhen Li, Edward Adelson, and Aseem Agarwala. “ScribbleBoost: Adding Classification to Edge-Aware Interpolation of Local Image and Video Adjustments”. In: Computer Graphics Forum. Vol. 27. 4. Wiley Online Library. 2008, pp. 1255–1264 (cit. on p. 33).

[LS18] Zhengqi Li and Noah Snavely. “Learning Intrinsic Image Decomposition From Watching the World”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). June 2018 (cit. on p. 78).

[Li+20] Zhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli, and Man-mohan Chandraker. “Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF From a Single Image”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 2475–2484 (cit. on pp. 72, 74, 87–91, 95, 99, 114, 203).

[LSC18] Zhengqin Li, Kalyan Sunkavalli, and Manmohan Chandraker. “Materials for Masses: SVBRDF Acquisition with a Single Mobile Phone Image”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 72–87 (cit. on pp. 72, 125).

[Li+18] Zhengqin Li, Zexiang Xu, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. “Learning to Reconstruct Shape and Spatially-varying Reflectance from a Single Image”. In: ACM Transactions on Graphics (TOG) 37.6 (2018), pp. 1–11 (cit. on pp. 125, 129).

[LLK19] Junbang Liang, Ming Lin, and Vladlen Koltun. “Differentiable Cloth Simulation for Inverse Problems”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on p. 158).

[Lia+17] Jing Liao, Yuan Yao, Lu Yuan, Gang Hua, and Sing Bing Kang. “Visual Attribute Transfer through Deep Image Analogy”. In: ACM Transactions on Graphics (TOG) 36.4 (2017), pp. 1–15 (cit. on pp. 31, 32, 46, 47, 56, 57, 60).

[Lic+21] Daniel Lichy, Jiaye Wu, Soumyadip Sengupta, and David W Jacobs. “Shape and Material Capture at Home”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 6123–6133 (cit. on p. 129).

[LPG19] Yiming Lin, Pieter Peers, and Abhijeet Ghosh. “On-Site Example-Based Material Appearance Acquisition”. In: Computer Graphics Forum. Vol. 38. 4. Wiley Online Library. 2019, pp. 15–25 (cit. on pp. 34, 105).

[Liu+20] Guilin Liu, Rohan Taori, Ting-Chun Wang, Zhiding Yu, Shiqiu Liu, Fitsum A Reda, Karan Sapra, Andrew Tao, and Bryan Catanzaro. “Transposer: Universal Texture Synthesis Using Feature Maps as Transposed Convolution Filter”. In: arXiv preprint arXiv:2007.07243 (2020) (cit. on pp. 72, 74, 75, 92).

[Liu+19a] Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. “On the Variance of the Adaptive Learning Rate and Beyond”. In: arXiv preprint arXiv:1908.03265 (2019) (cit. on p. 110).

[Liu+19b] Ming-Yu Liu, Xun Huang, Arun Mallya, Tero Karras, Timo Aila, Jaakko Lehtinen, and Jan Kautz. “Few-shot Unsupervised Image-to-Image Translation”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2019, pp. 10551–10560 (cit. on p. 37).

[Liu+21] Steven Liu, Xiuming Zhang, Zhoutong Zhang, Richard Zhang, Jun-Yan Zhu, and Bryan Russell. “Editing conditional radiance fields”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 5773–5783 (cit. on p. 116).

[Liu+22a] Yuan Liu, Sida Peng, Lingjie Liu, Qianqian Wang, Peng Wang, Christian Theobalt, Xiaowei Zhou, and Wenping Wang. “Neural Rays for Occlusion-aware Image-based Rendering”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 7824–7833 (cit. on p. 203).

[Liu+22b] Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. “A Convnet for the 2020s”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 11976–11986 (cit. on pp. 108, 109, 119).

[LSD15] Jonathan Long, Evan Shelhamer, and Trevor Darrell. “Fully Convolutional Networks for Semantic Segmentation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2015, pp. 3431–3440 (cit. on pp. 83, 94).

[LH17] Ilya Loshchilov and Frank Hutter. “Decoupled Weight Decay Regularization”. In: arXiv preprint arXiv:1711.05101 (2017) (cit. on p. 176).

[Lua+21] Fujun Luan, Shuang Zhao, Kavita Bala, and Zhao Dong. “Unified Shape and SVBRDF Recovery Using Differentiable Monte Carlo Rendering”. In: Computer Graphics Forum. Vol. 40. 4. Wiley Online Library. 2021, pp. 101–113 (cit. on pp. 125, 129).

[Lug+20] Andreas Lugmayr, Martin Danelljan, Luc Van Gool, and Radu Timofte. “Srflow: Learning the Super-resolution Space with Normalizing Flow”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2020, pp. 715–732 (cit. on p. 186).

[LM08] Christiane Luible and Nadia Magnenat-Thalmann. “The Simulation of Cloth Using Accurate Physical Parameters”. In: Proceedings of the Tenth IASTED International Conference on Computer Graphics and Imaging (CGIM ’08). 2008 (cit. on p. 158).

[LKS15] Zhaoliang Lun, Evangelos Kalogerakis, and Alla Sheffer. “Elements of Style: Learning Perceptual Shape Style Similarity”. In: ACM Transactions on graphics (TOG) 34.4 (2015), pp. 1–14 (cit. on p. 160).

[MHN13] Andrew L Maas, Awni Y Hannun, and Andrew Y Ng. “Rectifier Nonlinearities Improve Neural Network Acoustic Models”. In: International Conference on Machine Learning (ICML). Vol. 30. 1. Citeseer. 2013, p. 3 (cit. on pp. 77, 94, 176).

[Mag+07] Nadia Magnenat-Thalmann, Christiane Luible, Pascal Volino, and Etienne Lyard. “From Measured Fabric to the Simulation of Cloth”. In: 2007 10th IEEE International Conference on Computer-Aided Design and Computer Graphics. IEEE. 2007, pp. 7–18 (cit. on p. 158).

[MR10] Sébastien Marcel and Yann Rodriguez. “Torchvision the Machine-vision Package of Torch”. In: Proceedings of the 18th ACM International Conference on Multimedia. 2010, pp. 1485–1488 (cit. on pp. 67, 84, 110, 142, 194).

[Mar+20] Morteza Mardani, Guilin Liu, Aysegul Dundar, Shiqiu Liu, Andrew Tao, and Bryan Catanzaro. “Neural FFTs for Universal Texture Image Synthesis”. In: Advances in Neural Information Processing Systems 33 (2020) (cit. on pp. 72, 74, 75, 85, 107, 128).

[MFS08] Bruce A Maxwell, Richard M Friedhoff, and Casey A Smith. “A biilluminant dichromatic reflection model for understanding images”. In: 2008 IEEE Conference on Computer Vision and Pattern Recognition. IEEE. 2008, pp. 1–8 (cit. on p. 10).

[Maz+19] Ilya Mazlov, Sebastian Merzbach, Elena Trunz, and Reinhard Klein. “Neural Appearance Synthesis and Transfer”. In: (2019), pp. 35–39 (cit. on p. 31).

[MHM18] Leland McInnes, John Healy, and James Melville. “Umap: Uniform Manifold Approximation and Projection for Dimension Reduction”. In: arXiv preprint arXiv:1802.03426 (2018) (cit. on p. 133).

[Meh+21] Ishit Mehta, Michaël Gharbi, Connelly Barnes, Eli Shechtman, Ravi Ramamoorthi, and Manmohan Chandraker. “Modulated Periodic Activations for Generalizable Local Functional Representations”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 14214–14223 (cit. on pp. 109, 186, 203).

[MR21] Sachin Mehta and Mohammad Rastegari. “MobileViT: Light-Weight, General-Purpose, and Mobile-Friendly Vision Transformer”. In: International Conference on Learning Representations (ICLR). 2021 (cit. on pp. 124, 127, 132, 141, 144).

[Mel+12] Francho Melendez, Mashhuda Glencross, Jon Starck, and Gregory J Ward. “Transfer of Albedo and Local Depth Variation to Photo-Textures”. In: Proceedings of the 9th European Conference on Visual Media Production. 2012, pp. 40–48 (cit. on pp. 31, 34).

[Mel+21] Joe Mellor, Jack Turner, Amos Storkey, and Elliot J Crowley. “Neural Architecture Search Without Training”. In: International Conference on Machine Learning (ICML). PMLR. 2021, pp. 7588–7598 (cit. on p. 92).

[Mer+19] Sebastian Merzbach, Max Hermann, Martin Rump, and Reinhard Klein. “Learned Fitting of Spatially Varying BRDFs”. In: Computer Graphics Forum. Vol. 38. 4. Wiley Online Library. 2019, pp. 193–205 (cit. on p. 30).

[MWK17] Sebastian Merzbach, Michael Weinmann, and Reinhard Klein. “High-quality Multi-Spectral Reflectance Acquisition with X-rite TAC7”. In: Proceedings of the Workshop on Material Appearance Modeling. 2017, pp. 11–16 (cit. on pp. 35, 39).

[MNG17] Lars Mescheder, Sebastian Nowozin, and Andreas Geiger. “Adversarial Variational Bayes: Unifying Variational Autoencoders and Generative Adversarial Networks”. In: International Conference on Machine Learning (ICML). PMLR. 2017, pp. 2391–2400 (cit. on p. 126).

[Mic+18] Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. “Mixed Precision Training”. In: International Conference on Learning Representations (ICLR). 2018 (cit. on pp. 56, 67, 84, 94, 110, 142, 176, 194).

[Mig+12] Eder Miguel, Derek Bradley, Bernhard Thomaszewski, Bernd Bickel, Wojciech Matusik, Miguel A Otaduy, and Steve Marschner. “Data-driven Estimation of Cloth Simulation Models”. In: Computer Graphics Forum. Vol. 31. 2pt2. Wiley Online Library. 2012, pp. 519–528 (cit. on p. 158).

[Mil+20] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2020 (cit. on pp. 186, 187, 207, 212, 262, 264).

[Min95] Pier Giorgio Minazio. “FAST–Fabric Assurance by Simple Testing”. In: International Journal of Clothing Science and Technology (1995) (cit. on pp. 156, 158).

[Miy+18] Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. “Spectral Normalization for Generative Adversarial Networks”. In: International Conference on Learning Representations. 2018 (cit. on pp. 141, 145).

[Mor+20] Alexander Mordvintsev, Ettore Randazzo, Eyvind Niklasson, and Michael Levin. “Growing Neural Cellular Automata”. In: Distill (2020). https://distill.pub/2020/growing-ca (cit. on pp. 15, 75, 257).

[Mor15] Andrew Morgan. The True Cost. 2015 (cit. on p. 2).

[Mor+17] Joep Moritz, Stuart James, Tom S.F. Haines, Tobias Ritschel, and Tim Weyrich. “Texture Stationarization: Turning Photos into Tileable Textures”. In: Eurographics Symposium on Geometry Processing. Vol. 36. 2. 2017, pp. 177–188 (cit. on pp. 15, 72–74, 88–90, 99, 114, 257).

[MMK03] Gero Müller, Jan Meseth, and Reinhard Klein. “Compression and Real-Time Rendering of Measured BTFs Using Local PCA.” In: VMV. 2003, pp. 271–279 (cit. on p. 103).

[Mül+22] Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. “Instant Neural Graphics Primitives with a Multiresolution Hash Encoding”. In: ACM Transactions on Graphics (TOG) 41.4 (2022), 102:1–102:15 (cit. on pp. 203, 212).

[Mül+19] Thomas Müller, Brian McWilliams, Fabrice Rousselle, Markus Gross, and Jan Novák. “Neural Importance Sampling”. In: ACM Transactions on Graphics (TOG) 38.5 (2019), pp. 1–19 (cit. on pp. 116, 186, 191, 196, 197).

[Mül+20] Thomas Müller, Fabrice Rousselle, Alexander Keller, and Jan Novák. “Neural Control Variates”. In: ACM Transactions on Graphics (TOG) 39.6 (2020), pp. 1–19 (cit. on pp. 186, 191).

[Mun+22] Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas Müller, and Sanja Fidler. “Extracting Triangular 3D Models, Materials, and Lighting From Images”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 8280–8290 (cit. on pp. 127, 129).

[Mur+20] J Krishna Murthy, Miles Macklin, Florian Golemo, Vikram Voleti, Linda Petrini, Martin Weiss, Breandan Considine, Jérôme Parent-Lévesque, Kevin Xie, Kenny Erleben, et al. “GradSim: Differentiable simulation for System Identification and Visuomotor Control”. In: International Conference on Learning Representations (ICLR). 2020 (cit. on p. 158).

[Naf+16] Hossein Ziaei Nafchi, Atena Shahkolaei, Rachid Hedjam, and Mohamed Cheriet. “Mean Deviation Similarity Index: Efficient and Reliable Full-reference Image Quality Evaluator”. In: IEEE Access 4 (2016), pp. 5579–5590 (cit. on p. 160).

[Nag+15] Koki Nagano, Graham Fyffe, Oleg Alexander, Jernej Barbic, Hao Li, Abhijeet Ghosh, and Paul E Debevec. “Skin Microstructure Deformation with Displacement Map Convolution.” In: ACM Transactions on Graphics (TOG) 34.4 (2015), pp. 109–1 (cit. on p. 105).

[NH10] Vinod Nair and Geoffrey E Hinton. “Rectified Linear Units Improve Restricted Boltzmann Machines”. In: International Conference on Machine Learning (ICML). 2010 (cit. on pp. 54, 109, 111, 120).

[Nam+16] Giljoo Nam, Joo Ho Lee, Hongzhi Wu, Diego Gutierrez, and Min H Kim. “Simultaneous Acquisition of Microscale Reflectance and Normals”. In: ACM Transactions on Graphics (TOG) 35.6 (2016), pp. 1–11 (cit. on pp. 35, 39).

[NK18] Hyeonseob Nam and Hyo Eun Kim. “Batch-Instance Normalization for Adaptively Style-Invariant Neural Networks”. In: Advances in Neural Information Processing Systems. Vol. 2018-Decem. 2018, pp. 2558–2567 (cit. on p. 77).

[NSO12] Rahul Narain, Armin Samii, and James F O’brien. “Adaptive Anisotropic Remeshing for Cloth Simulation”. In: ACM Transactions on Graphics (TOG) 31.6 (2012), pp. 1–10 (cit. on p. 162).

[ND06] Addy Ngan and Frédo Durand. “Statistical Acquisition of Texture Appearance”. In: Proceedings of the 17th Eurographics Conference on Rendering Techniques. 2006, pp. 31–40 (cit. on p. 63).

[NJR15] Jannik Boll Nielsen, Henrik Wann Jensen, and Ravi Ramamoorthi. “On Optimal, Minimal BRDF Sampling for Reflectance Acquisition”. In: ACM Transactions on Graphics (TOG) 34.6 (2015), pp. 1–11 (cit. on pp. 42, 116, 130).

[Nik+21] Eyvind Niklasson, Alexander Mordvintsev, Ettore Randazzo, and Michael Levin. “Self-Organising Textures”. In: Distill (2021). https://distill.pub/selforg/2021/textures (cit. on pp. 72, 74, 75, 89–91, 93, 95, 99).

[Nim+19] Merlin Nimier-David, Delio Vicini, Tizian Zeltner, and Wenzel Jakob. “Mitsuba 2: A Retargetable Forward and Inverse Renderer”. In: ACM Transactions on Graphics (TOG) 38.6 (2019), pp. 1–17 (cit. on pp. 6, 109, 202).

[ODO16] Augustus Odena, Vincent Dumoulin, and Chris Olah. “Deconvolution and Checker-board Artifacts”. In: Distill (2016) (cit. on p. 109).

[Ouy+21] Yaobin Ouyang, Shiqiu Liu, Markus Kettunen, Matt Pharr, and Jacopo Pantaleoni. “ReSTIR GI: Path Resampling for Real-time Path Tracing”. In: Computer Graphics Forum. Vol. 40. 8. 2021, pp. 17–29 (cit. on p. 184).

[Pap+21] George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. “Normalizing Flows for Probabilistic Modeling and Inference”. In: The Journal of Machine Learning Research 22.1 (2021), pp. 2617–2680 (cit. on p. 191).

[PPM17] George Papamakarios, Theo Pavlakou, and Iain Murray. “Masked Autoregressive Flow for Density Estimation”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 191).

[PSM19] George Papamakarios, David Sterratt, and Iain Murray. “Sequential Neural Likelihood: Fast Likelihood-free Inference with Autoregressive Flows”. In: The 22nd International Conference on Artificial Intelligence and Statistics. PMLR. 2019, pp. 837–848 (cit. on p. 191).

[Par+20] Taesung Park, Alexei A Efros, Richard Zhang, and Jun-Yan Zhu. “Contrastive Learning for Unpaired Image-to-Image Translation”. In: arXiv preprint arXiv:2007.15651 (2020) (cit. on pp. 37, 210).

[Pas+19] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. “Pytorch: An Imperative Style, High-performance Deep Learning Library”. In: Advances in Neural Information Processing Systems. Vol. 32. 2019 (cit. on pp. 39, 55, 67, 84, 94, 110, 142, 176, 194).

[Pha+20] Minh Quang Pham, Josep-Maria Crego, François Yvon, and Jean Senellart. “A study of Residual Adapters for Multi-domain Neural Machine Translation”. In: Conference on Machine Translation. 2020 (cit. on p. 159).

[Pha19] Matt Pharr. Visualizing Warping Strategies for Sampling Environment Map Lights. 2019 (cit. on pp. 189, 199, 200).

[PJH20] Matt Pharr, Wenzel Jakob, and Greg Humphreys. Implementation of the forthcoming 4th edition of Physically Based Rendering: From Theory to Implementation. 2020 (cit. on pp. 3, 109, 189, 200).

[PJH16] Matt Pharr, Wenzel Jakob, and Greg Humphreys. Physically Based Rendering: From Theory to Implementation. 3rd. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 2016 (cit. on pp. 185, 189, 200, 201).

[PS00] Javier Portilla and Eero P Simoncelli. “A Parametric Texture Model Based on Joint Statistics of Complex Wavelet Coefficients”. In: International Journal of Computer Vision 40.1 (2000), pp. 49–70 (cit. on p. 74).

[Pow13] Jess Power. “Fabric Objective Measurements for Commercial 3D Virtual Garment Simulation”. In: International Journal of Clothing Science and Technology (2013) (cit. on p. 158).

[Pra+18] Ekta Prashnani, Hong Cai, Yasamin Mostofi, and Pradeep Sen. “Pieapp: Perceptual Image-error Assessment Through Pairwise Preference”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 1808–1817 (cit. on p. 160).

[Raa+18] Lara Raad, Axel Davy, Agnès Desolneux, and Jean-Michel Morel. “A Survey of Exemplar-Based Texture Synthesis”. In: Annals of Mathematical Sciences and Applications 3.1 (2018), pp. 89–148 (cit. on p. 73).

[Rad+21] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. “Learning Transferable Visual Models from Natural Language Supervision”. In: International Conference on Machine Learning (ICML). PMLR. 2021, pp. 8748–8763 (cit. on p. 159).

[Rag+19] Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio. “Transfusion: Under-standing Transfer Learning for Medical Imaging”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on p. 159).

[Rai+22] Gilles Rainer, Adrien Bousseau, Tobias Ritschel, and George Drettakis. “Neural Pre-computed Radiance Transfer”. In: Computer Graphics Forum. Vol. 41. 2. Wiley Online Library. 2022, pp. 365–378 (cit. on pp. 105, 116, 203).

[Rai+20] Gilles Rainer, Abhijeet Ghosh, Wenzel Jakob, and Tim Weyrich. “Unified Neural Encod-ing of BTFs”. In: Computer Graphics Forum. Vol. 39. 2. The Eurographics Association. 2020, pp. 1–13 (cit. on pp. 18, 33, 102, 103, 109, 112, 113, 116, 259).

[Rai+19] Gilles Rainer, Wenzel Jakob, Abhijeet Ghosh, and Tim Weyrich. “Neural BTF Compres-sion and Interpolation”. In: Computer Graphics Forum. Vol. 38. 2. Wiley Online Library. 2019, pp. 235–244 (cit. on pp. 18, 33, 39, 40, 102, 103, 106, 109, 112, 113, 259).

[Ram+19] Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. “Stand-alone Self-attention in Vision Models”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on p. 164).

[RH01a] Ravi Ramamoorthi and Pat Hanrahan. “An Efficient Representation for Irradiance Environment Maps”. In: Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques. SIGGRAPH ’01. New York, NY, USA: Association for Computing Machinery, 2001, pp. 497–500 (cit. on p. 13).

[RH01b] Ravi Ramamoorthi and Pat Hanrahan. “An Efficient Representation for Irradiance Environment Maps”. In: Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques. 2001, pp. 497–500 (cit. on pp. 196, 199, 202).

[Ras+20] Abdullah Haroon Rasheed, Victor Romero, Florence Bertails-Descoubes, Stefanie Wuhrer, Jean-Sébastien Franco, and Arnaud Lazarus. “Learning to Measure the Static Friction Coefficient in Cloth Contact”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 9912–9921 (cit. on p. 159).

[RD17] Ajita Rattani and Reza Derakhshani. “On Fine-tuning Convolutional Neural Networks for Smartphone Based Ocular Recognition”. In: 2017 IEEE international joint conference on biometrics (IJCB). IEEE. 2017, pp. 762–767 (cit. on p. 159).

[RBV18] Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. “Efficient Parametrization of Multi-domain Deep Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 8119–8127 (cit. on p. 159).

[RBV17] Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. “Learning Multiple Visual Domains with Residual Adapters”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 159).

[RJ19] A Sai Bharadwaj Reddy and D Sujitha Juliet. “Transfer Learning with ResNet-50 for Malaria Cell-image Classification”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). IEEE. 2019, pp. 0945–0949 (cit. on p. 159).

[Ree+15] Scott E Reed, Yi Zhang, Yuting Zhang, and Honglak Lee. “Deep Visual Analogy-Making”. In: Advances in Neural Information Processing Systems. 2015, pp. 1252–1260 (cit. on p. 32).

[Rei+14] Klein Reinhard, Rump Martin, Weinmann Michael, Sarlette Ralf, and Schwartz Christopher. “Design and Implementation of Practical Bidirectional Texture Function Measurement Devices focusing on the Developments at the University of Bonn”. In: Sensors 14.5 (2014) (cit. on pp. 109, 114, 117).

[Rei+18] Rafael Reisenhofer, Sebastian Bosse, Gitta Kutyniok, and Thomas Wiegand. “A Haar Wavelet-based Perceptual Similarity Index for Image Quality Assessment”. In: Signal Processing: Image Communication 61 (2018), pp. 33–43 (cit. on p. 160).

[RSS16] Nathalie Remy, Eveline Speelman, and Steven Swartz. Style that’s Sustainable: A New Fast-Fashion Formula. 2016 (cit. on p. 2).

[Ren+17] Jian Ren, Xiaohui Shen, Zhe Lin, Radomir Mech, and David J Foran. “Personalized Image Aesthetics”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2017, pp. 638–647 (cit. on p. 74).

[Rez+20] Danilo Jimenez Rezende, George Papamakarios, Sébastien Racaniere, Michael Albergo, Gurtej Kanwar, Phiala Shanahan, and Kyle Cranmer. “Normalizing Flows on Tori and Spheres”. In: International Conference on Machine Learning (ICML). PMLR. 2020, pp. 8083–8092 (cit. on p. 189).

[Al-+16] Rami Al-Rfou, Guillaume Alain, Amjad Almahairi, Christof Angermueller, Dzmitry Bahdanau, Nicolas Ballas, Frédéric Bastien, Justin Bayer, Anatoly Belikov, Alexander Belopolsky, et al. “Theano: A Python Framework for Fast Computation of Mathematical Expressions”. In: arXiv e-prints (2016), arXiv–1605 (cit. on p. 95).

[Rho+22] Daniel Rho, Junwoo Cho, Jong Hwan Ko, and Eunbyung Park. “Neural Residual Flow Fields for Efficient Video Representations”. In: Proceedings of the Asian Conference on Computer Vision. 2022, pp. 3447–3463 (cit. on p. 187).

[Rib+20] Edgar Riba, Dmytro Mishkin, Daniel Ponsa, Ethan Rublee, and Gary Bradski. “Kornia: An Open Source Differentiable Computer Vision Library for Pytorch”. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2020, pp. 3674–3683 (cit. on pp. 67, 111, 142, 176).

[Ric+16] Stephan R. Richter, Vibhav Vineet, Stefan Roth, and Vladlen Koltun. “Playing for Data: Ground Truth from Computer Games”. In: European Conference on Computer Vision (ECCV). Ed. by Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling. Vol. 9906. LNCS. Springer International Publishing, 2016, pp. 102–118 (cit. on p. 212).

[RPG16] Jérémy Riviere, Pieter Peers, and Abhijeet Ghosh. “Mobile Surface Reflectometry”. In: Computer Graphics Forum. Vol. 35. 1. Wiley Online Library. 2016, pp. 191–202 (cit. on pp. 31, 34).

[Rod+23a] Carlos Rodriguez-Pardo, Henar Dominguez-Elvira, David Pascual-Hernandez, and Elena Garces. “UMat: Uncertainty-Aware Single Image High Resolution Material Capture”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023) (cit. on pp. 24, 62, 76, 102, 107, 108, 110, 116, 193, 268).

[RG21] Carlos Rodriguez-Pardo and Elena Garces. “Neural Photometry-guided Visual Attribute Transfer”. In: IEEE Transactions on Visualization and Computer Graphics (2021) (cit. on pp. 23, 25, 63–70, 102, 105, 107, 108, 110, 114, 125, 129, 131, 142, 193, 268).

[RG22] Carlos Rodriguez-Pardo and Elena Garces. “SeamlessGAN: Self-Supervised Synthesis of Tileable Texture Maps”. In: IEEE Transactions on Visualization and Computer Graphics (2022) (cit. on pp. 23, 25, 64, 67, 69, 104, 107, 108, 110, 114, 127, 128, 141, 268).

[Rod+23b] Carlos Rodriguez-Pardo, Melania Prieto-Martin, Dan Casas, and Elena Garces. “How Will It Drape Like? Capturing Fabric Mechanics from Depth Images”. In: Computer Graphics Forum (Proc. of Eurographics) (2023) (cit. on pp. 18, 23, 212, 268).

[Rod+19] Carlos Rodriguez-Pardo, Sergio Suja, David Pascual, Jorge Lopez-Moreno, and Elena Garces. “Automatic Extraction and Synthesis of Regular Repeatable Patterns”. In: Computers & Graphics 83 (2019), pp. 33–41 (cit. on pp. 15, 34, 72, 75, 79, 83, 87, 89–91, 93, 95, 99, 114, 257).

[RB19] Carlos Rodrıguez-Pardo and Hakan Bilen. “Personalised Aesthetics with Residual Adapters”. In: Iberian Conference on Pattern Recognition and Image Analysis. Springer. 2019, pp. 508–520 (cit. on pp. 74, 90, 159).

[Rom+22] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. “High-Resolution Image Synthesis with Latent Diffusion Models”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 10684–10695 (cit. on pp. 127, 209, 210).

[RFB15] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. “U-Net: Convolutional Networks for Biomedical Image Segmentation”. In: International Conference on Medical image computing and computer-assisted intervention. Springer. 2015, pp. 234–241 (cit. on pp. 37–39, 54, 55, 64, 66, 67, 69, 78, 108, 111, 124, 127, 132, 141).

[RSK13] Roland Ruiters, Christopher Schwartz, and Reinhard Klein. “Example-based Interpolation and Synthesis of Bidirectional Texture Functions”. In: Computer Graphics Forum. Vol. 32. 2pt3. Wiley Online Library. 2013, pp. 361–370 (cit. on p. 63).

[RSK10] Martin Rump, Ralf Sarlette, and Reinhard Klein. “Groundtruth Data for Multispectral Bidirectional Texture Functions”. In: Conference on Colour in Graphics, Imaging, and Vision. Vol. 2010. 1. Society for Imaging Science and Technology. 2010, pp. 326–331 (cit. on p. 116).

[Run+20] Tom FH Runia, Kirill Gavrilyuk, Cees GM Snoek, and Arnold WM Smeulders. “Cloth in the Wind: A Case Study of Physical Measurement Through Simulation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 10498–10507 (cit. on p. 158).

[Rus+04] Daniel B Russakoff, Carlo Tomasi, Torsten Rohlfing, and Calvin R Maurer. “Image Similarity Using Mutual Information of Regions”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2004, pp. 596–607 (cit. on p. 132).

[Sah+22] Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. “Palette: Image-to-Image Diffusion Models”. In: ACM SIGGRAPH 2022 Conference Proceedings. 2022, pp. 1–10 (cit. on pp. 127, 209).

[San+19] Veit Sandfort, Ke Yan, Perry J Pickhardt, and Ronald M Summers. “Data Augmentation Using Generative Adversarial Networks (CycleGAN) to Improve Generalizability in CT Segmentation Tasks”. In: Scientific reports 9.1 (2019), pp. 1–9 (cit. on p. 45).

[SC20] Shen Sang and Manmohan Chandraker. “Single-Shot Neural Relighting and SVBRDF Estimation”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2020, pp. 85–101 (cit. on p. 125).

[SOC22] Igor Santesteban, Miguel A Otaduy, and Dan Casas. “SNUG: Self-Supervised Neural Dynamic Garments”. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022) (cit. on pp. 17, 259).

[SSK03] Mirko Sattler, Ralf Sarlette, and Reinhard Klein. “Efficient and Realistic Visualization of Cloth”. In: Rendering techniques. 2003, pp. 167–178 (cit. on p. 109).

[SSK20] Edgar Schonfeld, Bernt Schiele, and Anna Khoreva. “A U-Net Based Discriminator for Generative Adversarial Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 8207–8216 (cit. on pp. 83, 93, 124, 128, 132, 141, 142).

[Sch+17] Vincent Schüssler, Eric Heitz, Johannes Hanika, and Carsten Dachsbacher. “Microfacet-based Normal Mapping for Robust Monte Carlo Path Tracing”. In: ACM Transactions on Graphics (TOG) 36.6 (2017), pp. 1–12 (cit. on p. 201).

[See66] Robert T Seeley. “Spherical Harmonics”. In: The American Mathematical Monthly 73.4P2 (1966), pp. 115–121 (cit. on pp. 185, 196).

[Sel+20] Raghavendra Selvan, Frederik Faye, Jon Middleton, and Akshay Pai. “Uncertainty Quantification in Medical Image Segmentation with Normalizing Flows”. In: Machine Learning in Medical Imaging. 2020, pp. 80–90 (cit. on p. 186).

[SKK18] Murat Sensoy, Lance Kaplan, and Melih Kandemir. “Evidential Deep Learning to Quantify Classification Uncertainty”. In: Advances in Neural Information Processing Systems 31 (2018) (cit. on pp. 126, 130).

[SDM19] Tamar Rott Shaham, Tali Dekel, and Tomer Michaeli. “SinGAN: Learning a Generative Model From a Single Natural Image”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2019 (cit. on pp. 32, 38, 57, 75, 85, 86, 90).

[Shi+20] Liang Shi, Beichen Li, Miloš Hašan, Kalyan Sunkavalli, Tamy Boubekeur, Radomir Mech, and Wojciech Matusik. “Match: Differentiable Material Graphs for Procedural Material Capture”. In: ACM Transactions on Graphics (TOG) (2020) (cit. on pp. 76, 125, 137, 138, 143, 151–153).

[Sho+19] Assaf Shocher, Shai Bagon, Phillip Isola, and Michal Irani. “InGAN: Capturing and Retargeting the "DNA" of a Natural Image”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2019 (cit. on p. 75).

[SCI18] Assaf Shocher, Nadav Cohen, and Michal Irani. ““Zero-Shot” Super-Resolution Using Deep Internal Learning”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2018 (cit. on p. 75).

[SK19] Connor Shorten and Taghi M Khoshgoftaar. “A Survey on Image Data Augmentation for Deep Learning”. In: Journal of Big Data 6.1 (2019), p. 60 (cit. on pp. 36, 45).

[SZ15] Karen Simonyan and Andrew Zisserman. “Very Deep Convolutional Networks for Large-Scale Image Recognition”. In: International Conference on Learning Representations (ICLR). 2015 (cit. on pp. 32, 164).

[Sin+21] Abhishek Sinha, Kumar Ayush, Jiaming Song, Burak Uzkent, Hongxia Jin, and Stefano Ermon. “Negative Data Augmentation”. In: arXiv preprint arXiv:2102.05113 (2021) (cit. on p. 93).

[Sit+20] Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. “Implicit Neural Representations with Periodic Activation Functions”. In: Advances in Neural Information Processing Systems 33 (2020) (cit. on pp. 75, 102, 109, 112, 120, 186, 193, 194, 262).

[Sit+21] Vincent Sitzmann, Semon Rezchikov, William T. Freeman, Joshua B. Tenenbaum, and Fredo Durand. “Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering”. In: Advances in Neural Information Processing Systems. 2021 (cit. on p. 187).

[SKS02] Peter-Pike Sloan, Jan Kautz, and John Snyder. “Precomputed Radiance Transfer for Real-Time Rendering in Dynamic, Low-Frequency Lighting Environments”. In: ACM Transactions on Graphics (TOG) 21.3 (July 2002), pp. 527–536 (cit. on p. 13).

[Sne17] Xavier Snelgrove. “High-resolution Multi-scale Neural Texture Synthesis”. In: SIG-GRAPH Asia 2017 Technical Briefs. 2017, pp. 1–4 (cit. on p. 74).

[Sol+21] Ava P Soleimany, Alexander Amini, Samuel Goldman, Daniela Rus, Sangeeta N Bhatia, and Connor W Coley. “Evidential Deep Learning for Guided Molecular Property Prediction and Discovery”. In: ACS Central Science 7.8 (2021), pp. 1356–1367 (cit. on pp. 126, 135).

[Spe+22] Georg Sperl, Rosa M Sánchez-Banderas, Manwen Li, Chris Wojtan, and Miguel A Otaduy. “Estimation of Yarn-level Simulation Models for Production Fabrics”. In: ACM Transactions on Graphics (TOG) 41.4 (2022), pp. 1–15 (cit. on p. 163).

[Sri+20] Pratul P Srinivasan, Ben Mildenhall, Matthew Tancik, Jonathan T Barron, Richard Tucker, and Noah Snavely. “Lighthouse: Predicting Lighting Volumes for Spatially-Coherent Illumination”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 8080–8089 (cit. on p. 203).

[Sri+14] Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”. In: The Journal of Machine Learning Research 15.1 (2014), pp. 1929–1958 (cit. on pp. 110, 127, 141, 176).

[Ste+14] Heinz Christian Steinhausen, Dennis den Brok, Matthias B Hullin, and Reinhard Klein. “Acquiring Bidirectional Texture Functions for Large-Scale Material Samples”. In: (2014) (cit. on pp. 30, 63).

[Ste+15a] Heinz Christian Steinhausen, Dennis den Brok, Matthias B. Hullin, and Reinhard Klein. “Extrapolating Large-Scale Material BTFs under Cross-Device Constraints”. In: Vision, Modeling & Visualization. Ed. by David Bommes, Tobias Ritschel, and Thomas Schultz. The Eurographics Association, 2015, pp. 143–150 (cit. on pp. 33, 104).

[Ste+15b] Heinz Christian Steinhausen, Rodrigo Martín, Dennis den Brok, Matthias B. Hullin, and Reinhard Klein. “Extrapolation of Bidirectional Texture Functions using Texture Synthesis guided by Photometric Normals”. In: Measuring, Modeling, and Reproducing Material Appearance II (SPIE 9398). Vol. 9398. 14. San Francisco, USA, Feb. 2015 (cit. on pp. 33, 104).

[Str+22] Yannick Strümpler, Janis Postels, Ren Yang, Luc Van Gool, and Federico Tombari. “Implicit Neural Representations for Image Compression”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2022, pp. 74–91 (cit. on p. 187).

[SGK21] Vadim Sushko, Juergen Gall, and Anna Khoreva. “One-Shot GAN: Learning to Generate Samples from Single Images and Videos”. In: arXiv preprint arXiv:2103.13389 (2021) (cit. on p. 75).

[SB08] Cédric Syllebranque and Samuel Boivin. “Estimation of Mechanical Parameters of Deformable Solids from Videos”. In: The Visual Computer 24.11 (2008), pp. 963–972 (cit. on p. 158).

[Szt+21] Alejandro Sztrajman, Gilles Rainer, Tobias Ritschel, and Tim Weyrich. “Neural BRDF Representation and Importance Sampling”. In: Computer Graphics Forum. Vol. 40. 6. Wiley Online Library. 2021, pp. 332–346 (cit. on pp. 105, 111, 116, 186, 187, 193, 203, 210).

[Tak+21] Towaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis, Charles Loop, Derek Nowrouzezahrai, Alec Jacobson, Morgan McGuire, and Sanja Fidler. “Neural Geometric Level of Detail: Real-time Rendering with Implicit 3D Shapes”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021 (cit. on p. 186).

[Tam+11] Omer Tamuz, Ce Liu, Serge Belongie, Ohad Shamir, and Adam Tauman Kalai. “Adaptively Learning the Crowd Kernel”. In: International Conference on Machine Learning (ICML). 2011, pp. 673–680 (cit. on pp. 170, 171).

[Tan+22] Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul Srinivasan, Jonathan T. Barron, and Henrik Kretzschmar. “Block-NeRF: Scalable Large Scene Neural View Synthesis”. In: arXiv. 2022 (cit. on p. 187).

[Tan+20] Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. “Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains”. In: Advances in Neural Information Processing Systems (2020) (cit. on pp. 75, 186, 193).

[Tew+22] Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, Wang Yifan, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, et al. “Advances in Neural Rendering”. In: Computer Graphics Forum. Vol. 41. 2. Wiley Online Library. 2022, pp. 703–735 (cit. on pp. 187, 212).

[Tex+20a] Ondřej Texler, David Futschik, Jakub Fišer, Michal Lukáč, Jingwan Lu, Eli Shechtman, and Daniel Sy`kora. “Arbitrary Style Transfer using Neurally-Guided Patch-Based Ssynthesis”. In: Computers & Graphics 87 (2020), pp. 62–71 (cit. on pp. 39, 45).

[Tex+20b] Ondřej Texler, David Futschik, Michal Kučera, Ondřej Jamriška, Šárka Sochorová, Menclei Chai, Sergey Tulyakov, and Daniel Sy`kora. “Interactive Video Stylization Using Few-Shot Patch-Based Training”. In: ACM Transactions on Graphics (TOG) 39.4 (2020), pp. 73–1 (cit. on pp. 30, 31, 33, 37, 39, 46, 53, 58, 61, 107, 108, 128, 129).

[TOB15] Lucas Theis, Aäron van den Oord, and Matthias Bethge. “A Note on the Evaluation of Generative Models”. In: arXiv preprint arXiv:1511.01844 (2015) (cit. on p. 41).

[Tom94] Shoji Tominaga. “Dichromatic reflection models for a variety of materials”. In: Color Research & Application 19.4 (1994), pp. 277–285 (cit. on p. 10).

[Ton+02] Xin Tong, Jingdan Zhang, Ligang Liu, Xi Wang, Baining Guo, and Heung-Yeung Shum. “Synthesis of Bidirectional Texture Functions on Arbitrary Surfaces”. In: ACM Transactions on Graphics (TOG) 21.3 (2002), pp. 665–672 (cit. on p. 103).

[TS67] Kenneth E Torrance and Ephraim M Sparrow. “Theory for Off-Specular Reflection from Roughened Surfaces”. In: Josa 57.9 (1967), pp. 1105–1114 (cit. on p. 11).

[Tu+20] Peihan Tu, Li-Yi Wei, Koji Yatani, Takeo Igarashi, and Matthias Zwicker. “Continuous Curve Textures”. In: ACM Transactions on Graphics (TOG) 39.6 (2020), pp. 1–16 (cit. on p. 72).

[TK74] Amos Tversky and Daniel Kahneman. “Judgment under Uncertainty: Heuristics and Biases: Biases in Judgments Reveal Some Heuristics of Thinking Under Uncertainty.” In: Science 185.4157 (1974), pp. 1124–1131 (cit. on p. 169).

[Uly+16] Dmitry Ulyanov, Vadim Lebedev, Andrea Vedaldi, and Victor S Lempitsky. “Texture Networks: Feed-Forward Synthesis of Textures and Stylized Images.” In: International Conference of Machine Learning (ICML). Vol. 1. 2. 2016, p. 4 (cit. on p. 76).

[UVL18] Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. “Deep Image Prior”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2018 (cit. on p. 75).

[UVL16] Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. “Instance Normalization: The Missing Ingredient for Fast Stylization”. In: arXiv preprint arXiv:1607.08022 (2016) (cit. on pp. 77, 94).

[Van11] Dietger G Van Antwerpen. “Unbiased Physically Nased Rendering on the GPU”. MA thesis. Electrical Engineering, Mathematics and Computer Science, 2011 (cit. on p. 195).

[Vas+17] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. “Attention Is All You Need”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 164).

[VG95] Eric Veach and Leonidas J. Guibas. “Optimally Combining Sampling Techniques for Monte Carlo Rendering”. In: Proceedings of the 22nd Annual Conference on Computer Graphics and Interactive Techniques. SIGGRAPH ’95. Association for Computing Machinery, 1995 (cit. on p. 184).

[VPS21] Giuseppe Vecchio, Simone Palazzo, and Concetto Spampinato. “SurfaceNet: Adversarial SVBRDF Estimation from a Single Image”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 12840–12848 (cit. on pp. 17, 76, 123, 125, 128, 129, 132, 258).

[Ver+22] Dor Verbin, Peter Hedman, Ben Mildenhall, Todd Zickler, Jonathan T. Barron, and Pratul P. Srinivasan. “Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2022 (cit. on pp. 187, 212).

[VZH20] Yael Vinker, Nir Zabari, and Yedid Hoshen. “Training End-to-end Single Image Generators without GANs”. In: arXiv preprint arXiv:2004.06014 (2020) (cit. on p. 75).

[VMF09] Pascal Volino, Nadia Magnenat-Thalmann, and Francois Faure. “A Simple Approach to Nonlinear Tensile Stiffness for Accurate Cloth Simulation”. In: ACM Transactions on Graphics (TOG) 28.4 (2009), Article–No (cit. on pp. 158, 161, 162).

[Wal+14] Ingo Wald, Sven Woop, Carsten Benthin, Gregory S. Johnson, and Manfred Ernst. “Embree: A Kernel Framework for Efficient CPU Ray Tracing”. In: ACM Transactions on Graphics (TOG) 33.4 (2014) (cit. on p. 195).

[Wal+07] Bruce Walter, Stephen R Marschner, Hongsong Li, and Kenneth E Torrance. “Microfacet Models for Refraction through Rough Surfaces”. In: Proceedings of the 18th Eurographics Conference on Rendering Techniques. 2007, pp. 195–206 (cit. on p. 127).

[WZY22] Cairong Wang, Yiming Zhu, and Chun Yuan. “Diverse Image Inpainting with Normalizing Flow”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2022, pp. 53–69 (cit. on p. 186).

[Wan+22a] Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. “Clip-nerf: Text-and-image Driven Manipulation of Neural Radiance Fields”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 3835–3844 (cit. on p. 116).

[Wan+22b] Chen Wang, Xiang Wang, Jiawei Zhang, Liang Zhang, Xiao Bai, Xin Ning, Jun Zhou, and Edwin Hancock. “Uncertainty Estimation for Stereo Matching Based on Evidential Deep Learning”. In: Pattern Recognition 124 (2022), p. 108498 (cit. on p. 130).

[Wan+18] Guotai Wang, Wenqi Li, Maria A Zuluaga, Rosalind Pratt, Premal A Patel, Michael Aertsen, Tom Doel, Anna L David, Jan Deprest, Sébastien Ourselin, et al. “Interactive Medical Image Segmentation using Deep Learning with Image-specific Fine Tuning”. In: IEEE transactions on medical imaging 37.7 (2018), pp. 1562–1573 (cit. on p. 159).

[WOR11] Huamin Wang, James F O’Brien, and Ravi Ramamoorthi. “Data-driven Elastic Models for Cloth: Modeling and Measurement”. In: ACM Transactions on Graphics (TOG) 30.4 (2011), pp. 1–12 (cit. on p. 158).

[Wan+09] Jiaping Wang, Peiran Ren, Minmin Gong, John Snyder, and Baining Guo. “All-Frequency Rendering of Dynamic, Spatially-Varying Reflectance”. In: 28.5 (Dec. 2009), pp. 1–10 (cit. on p. 13).

[Wan+22c] Jiayi Wang, Diogo Luvizon, Franziska Mueller, Florian Bernard, Adam Kortylewski, Dan Casas, and Christian Theobalt. “HandFlow: Quantifying View-Dependent 3D Ambiguity in Two-Hand Reconstruction with Normalizing Flow”. In: International Symposium on Vision, Modeling, and Visualization. 2022 (cit. on p. 186).

[Wan+20a] Jiayun Wang, Yubei Chen, Rudrasis Chakraborty, and Stella X. Yu. “Orthogonal Convolutional Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2020 (cit. on p. 93).

[Wan+22d] Liao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu, and Lan Xu. “Fourier Plenoctrees for Dynamic Radiance Field Rendering in Real-time”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 13524–13534 (cit. on pp. 203, 212).

[Wan+20b] Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. “Linformer: Self-attention with Linear Complexity”. In: arXiv preprint arXiv:2006.04768 (2020) (cit. on pp. 124, 127, 141, 144).

[Wan+19] Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Bryan Catanzaro, and Jan Kautz. “Few-shot Video-to-Video Synthesis”. In: Advances in Neural Information Processing Systems. 2019, pp. 5013–5024 (cit. on p. 37).

[Wan+20c] Xiaofei Wang, Yiwen Han, Victor CM Leung, Dusit Niyato, Xueqiang Yan, and Xu Chen. “Convergence of Edge Computing and Deep Learning: A Comprehensive Survey”. In: IEEE Communications Surveys & Tutorials 22.2 (2020), pp. 869–904 (cit. on p. 212).

[Wan+20d] Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni. “Generalizing From a Few Examples: A Survey on Few-Shot Learning”. In: ACM Computing Surveys (CSUR) 53.3 (2020), pp. 1–34 (cit. on p. 37).

[Wan+04] Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. “Image Quality Assessment: From Error Visibility to Structural Similarity”. In: IEEE Transactions on Image Processing 13.4 (2004), pp. 600–612 (cit. on pp. 50, 54, 85, 86, 90, 92, 160, 198, 199, 202).

[Wan+21] Zian Wang, Jonah Philion, Sanja Fidler, and Jan Kautz. “Learning Indoor Inverse Rendering with 3D Spatially-Varying Lighting”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 12538–12547 (cit. on p. 203).

[Wan+23] Zian Wang, Tianchang Shen, Jun Gao, Shengyu Huang, Jacob Munkberg, Jon Hasselgren, Zan Gojcic, Wenzheng Chen, and Sanja Fidler. “Neural Fields meet Explicit Geometric Representations for Inverse Rendering of Urban Scenes”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2023 (cit. on pp. 203, 212).

[WGK14] Michael Weinmann, Juergen Gall, and Reinhard Klein. “Material Classification Based on Training Data Synthesized Using a BTF Database”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer International Publishing, 2014, pp. 156–171 (cit. on pp. 50, 103, 108, 111, 151).

[WKW16] Karl Weiss, Taghi M Khoshgoftaar, and DingDing Wang. “A Survey of Transfer Learning”. In: Journal of Big data 3.1 (2016), pp. 1–40 (cit. on p. 159).

[Wen+22] Tao Wen, Beibei Wang, Lei Zhang, Jie Guo, and Nicolas Holzschuch. “SVBRDF Recovery from a Single Image with Highlights Using a Pre-trained Generative Adversarial Network”. In: Computer Graphics Forum. Wiley Online Library. 2022 (cit. on p. 125).

[Woo+18] Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. “Cbam: Convolutional block attention module”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 3–19 (cit. on pp. 109, 119, 128, 141, 145, 164).

[WTX22] Qiling Wu, Jianchao Tan, and Kun Xu. “PaletteNeRF: Palette-based Color Editing for NeRFs”. In: arXiv preprint arXiv:2212.12871 (2022) (cit. on p. 116).

[WH18] Yuxin Wu and Kaiming He. “Group Normalization”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 3–19 (cit. on pp. 64, 67, 69, 127, 141, 144).

[Xie+22] Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, and Srinath Sridhar. “Neural Fields in Visual Computing and Beyond”. In: Computer Graphics Forum (2022) (cit. on pp. 187, 212).

[Xu+13] Kun Xu, Wei-Lun Sun, Zhao Dong, Dan-Yong Zhao, Run-Dong Wu, and Shi-Min Hu. “Anisotropic Spherical Gaussians”. In: ACM Transactions on Graphics (TOG) 32.6 (2013), pp. 1–11 (cit. on pp. 185, 196, 199, 202).

[Yan+21] Guandao Yang, Serge Belongie, Bharath Hariharan, and Vladlen Koltun. “Geometry processing with Neural Fields”. In: Advances in Neural Information Processing Systems. Vol. 34. 2021, pp. 22483–22497 (cit. on p. 186).

[YLL17] Shan Yang, Junbang Liang, and Ming C Lin. “Learning-based Cloth Material Recovery from Video”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2017, pp. 4383–4393 (cit. on pp. 156, 159).

[YL15] Shan Yang and Ming C Lin. “Materialcloning: Acquiring Elasticity Parameters from Images for Medical Applications”. In: IEEE Transactions on Visualization and Computer Graphics 22.9 (2015), pp. 2122–2135 (cit. on p. 158).

[Yan+18] Shan Yang, Zherong Pan, Tanya Amert, Ke Wang, Licheng Yu, Tamara Berg, and Ming C Lin. “Physics-inspired Garment Recovery from a Single-view Image”. In: ACM Transactions on Graphics (TOG) 37.5 (2018), pp. 1–14 (cit. on p. 158).

[Yao+22] Yao Yao, Jingyang Zhang, Jingbo Liu, Yihang Qu, Tian Fang, David McKinnon, Yanghai Tsin, and Long Quan. “NeILF: Neural Incident Light Field for Material and Lighting Estimation”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2022 (cit. on p. 187).

[Ye+22] Weicai Ye, Shuo Chen, Chong Bao, Hujun Bao, Marc Pollefeys, Zhaopeng Cui, and Guofeng Zhang. “Intrinsicnerf: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis”. In: arXiv preprint arXiv:2210.00647 (2022) (cit. on p. 116).

[Ye+21] Wenjie Ye, Yue Dong, Pieter Peers, and Baining Guo. “Deep Reflectance Scanning: Recovering Spatially-varying Material Appearance from a Flash-lit Video Sequence”. In: Computer Graphics Forum. Vol. 40. 6. Wiley Online Library. 2021, pp. 409–427 (cit. on p. 125).

[Ye+18] Wenjie Ye, Xiao Li, Yue Dong, Pieter Peers, and Xin Tong. “Single Image Surface Appearance Modeling with Self-Augmented CNNs and Inexact Supervision”. In: Computer Graphics Forum. Vol. 37. 7. Wiley Online Library. 2018, pp. 201–211 (cit. on pp. 34, 125).

[Yep20] Tyler Yep. torchinfo. Mar. 2020 (cit. on p. 113).

[YS19] Ye Yu and William AP Smith. “InverseRenderNet: Learning Single Image Inverse Rendering”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 3155–3164 (cit. on pp. 78, 203).

[YS21] Ye Yu and William AP Smith. “Outdoor Inverse Rendering from a Single Image using Multiview Self-supervision”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 44.7 (2021), pp. 3659–3675 (cit. on p. 203).

[Yun+19] Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. “Cutmix: Regularization strategy to train strong classifiers with localizable features”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 6023–6032 (cit. on p. 93).

[Zha+19a] Bo Zhang, Mingming He, Jing Liao, Pedro V Sander, Lu Yuan, Amine Bermak, and Dong Chen. “Deep Exemplar-based Video Colorization”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 8052–8061 (cit. on p. 33).

[Zha+19b] Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena. “Self-attention Generative Adversarial Networks”. In: International Conference on Machine Learning (ICML). PMLR. 2019, pp. 7354–7363 (cit. on pp. 164, 166, 176).

[ZDN16] Hang Zhang, Kristin Dana, and Ko Nishino. “Friction from Reflectance: Deep Reflectance Codes for Predicting Physical Surface Properties from One-shot In-field Reflectance”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2016, pp. 808–824 (cit. on p. 159).

[Zha+20] Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie. “Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity”. In: International Conference on Learning Representations (ICLR). 2020 (cit. on pp. 110, 194).

[Zha+21] Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely. “Physg: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 5453–5462 (cit. on p. 203).

[ZL12] Lin Zhang and Hongyu Li. “SR-SIM: A Fast and High Performance IQA Index Based on Spectral Residual”. In: IEEE International Conference on Image Processing. IEEE. 2012, pp. 1473–1476 (cit. on p. 160).

[ZSL14] Lin Zhang, Ying Shen, and Hongyu Li. “VSI: A Visual Saliency-induced Index for Perceptual Image Quality Assessment”. In: IEEE Transactions on Image processing 23.10 (2014), pp. 4270–4281 (cit. on p. 160).

[Zha+11] Lin Zhang, Lei Zhang, Xuanqin Mou, and David Zhang. “FSIM: A Feature Similarity Index for Image Quality Assessment”. In: IEEE Transactions on Image Processing 20.8 (2011), pp. 2378–2386 (cit. on p. 160).

[Zha+18] Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 586–595 (cit. on pp. 32, 39, 41, 49, 50, 54, 57, 67, 85, 86, 90, 110, 128, 142, 157, 160, 169, 170, 178, 198, 199, 202).

[ZJK20] Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. “Exploring Self-attention for Image Recognition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 10076–10085 (cit. on p. 164).

[Zho+20] Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. “Random Erasing Data Augmentation”. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 2020 (cit. on pp. 111, 129, 142, 176).

[Zho+16] Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. “Learning Deep Features for Discriminative Localization”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2016, pp. 2921–2929 (cit. on p. 164).

[Zho+22] Xilong Zhou, Miloš Hašan, Valentin Deschaintre, Paul Guerrero, Kalyan Sunkavalli, and Nima Kalantari. “TileGen: Tileable, Controllable Material Generation and Capture”. In: ACM Transactions on Graphics (Proc. SIGGRAPH Asia) (2022) (cit. on pp. 105, 107, 109, 116, 123).

[Zho+23] Xilong Zhou, Miloš Hašan, Valentin Deschaintre, Paul Guerrero, Kalyan Sunkavalli, and Nima Khademi Kalantari. “A Semi-Procedural Convolutional Material Prior”. In: Computer Graphics Forum. Wiley Online Library. 2023 (cit. on p. 105).

[ZK21] Xilong Zhou and Nima Khademi Kalantari. “Adversarial Single-Image SVBRDF Estimation with Hybrid Training”. In: Computer Graphics Forum. Vol. 40. 2. Wiley Online Library. 2021, pp. 315–325 (cit. on pp. 123, 125, 127, 137, 138, 141, 152, 153).

[Zho+17] Yang Zhou, Huajie Shi, Dani Lischinski, Minglun Gong, Johannes Kopf, and Hui Huang. “Analysis and Controlled Synthesis of Inhomogeneous Textures”. In: Computer Graphics Forum. Vol. 36. 2. 2017, pp. 199–212 (cit. on p. 74).

[Zho+18] Yang Zhou, Zhen Zhu, Xiang Bai, Dani Lischinski, Daniel Cohen-Or, and Hui Huang. “Non-Stationary Texture Synthesis by Adversarial Expansion”. In: ACM Transactions on Graphics (TOG) 37.4 (July 2018) (cit. on pp. 34, 38, 56, 72, 74, 76, 77, 79, 84, 85, 87, 90, 94, 104).

[Zho+19] Yizhou Zhou, Xiaoyan Sun, Zheng-Jun Zha, and Wenjun Zeng. “Context-Reinforced Semantic Segmentation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4046–4055 (cit. on p. 41).

[Zhu+17a] Jun Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. “Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2017-Octob (Mar. 2017), pp. 2242–2251 (cit. on pp. 78, 79, 83, 94).

[Zhu+17b] Jun-Yan Zhu, Richard Zhang, Deepak Pathak, Trevor Darrell, Alexei A Efros, Oliver Wang, and Eli Shechtman. “Toward Multimodal Image-to-Image Translation”. In: Advances in Neural Information Processing Systems. 2017, pp. 465–476 (cit. on p. 37).

[Zhu+21] Wei Zhu, Xian Guo, Dai Owaki, Kyo Kutsuzawa, and Mitsuhiro Hayashibe. “A Survey of Sim-to-real Transfer Techniques Applied to Reinforcement Learning for Bioinspired Robots”. In: IEEE Transactions on Neural Networks and Learning Systems (2021) (cit. on pp. 156, 212).

_______________________________

1 Additional publication details and results are included on the project website