CHAPTER 2. VISUAL ATTRIBUTE TRANSFER WITH PHOTOMETRY-GUIDED NEURAL NETWORKS

Figure 2.1: Illustration of the capabilities of our attribute transfer framework. On the left, we show a subset of our photometric dataset, comprised of several images of a material captured with different illumination conditions. On their right, two visual attributes: surface normals and segmentation maps. Using a larger image of the same or similar materials as guidance, our method can propagate the visual attributes, achieving high resolution, robust and controllable transfers at small computational cost.

Figure 2.1: Illustration of the capabilities of our attribute transfer framework. On the left, we show a subset of our photometric dataset, comprised of several images of a material captured with different illumination conditions. On their right, two visual attributes: surface normals and segmentation maps. Using a larger image of the same or similar materials as guidance, our method can propagate the visual attributes, achieving high resolution, robust and controllable transfers at small computational cost.

This chapter presents a deep learning-based method capable of propagating spatially-varying visual attributes of a material (e.g. texture maps or stylizations) to larger samples of the same or similar materials. By leveraging the material’s photometric response and a comprehensive data augmentation policy, we make the transfer robust to novel illumination conditions and affine deformations of the underlying surface. Our model relies on a supervised image-to-image translation framework and is agnostic to the transferred visual domain; we showcase a semantic segmentation, a normal map, and a stylization. Following an image analogies approach, the method only requires the training data to contain the same visual structures as the input guidance. Our approach works at interactive rates, making it suitable for material edit applications. We thoroughly evaluate our learning methodology in a controlled setup providing quantitative measures of performance. Last, we demonstrate that training the model on a single material is enough to generalize to materials of the same type without the need for massive datasets. The contributions presented in this chapter have led to the following publication1, which was also presented at the 2022 ACM SIGGRAPH Symposium on Interactive 3D Graphics and Games (I3D 2022) as an invited talk:

“Neural Photometry-guided Visual Attribute Transfer”
Carlos Rodriguez-Pardo, Elena Garces
IEEE Transactions on Visualization and Computer Graphics (TVCG)
(2021)

2.1 Introduction

Figure 2.2: The input exemplars on the left-bottom are transferred to the input guidances (second column) using images of the material taken under multiple illuminations (photometric input) as training data. Using this data for training, Texler et al. [Tex+20b] fail to generalize to geometric distortions and illumination conditions not present in the training set.

Figure 2.2: The input exemplars on the left-bottom are transferred to the input guidances (second column) using images of the material taken under multiple illuminations (photometric input) as training data. Using this data for training, Texler et al. [Tex+20b] fail to generalize to geometric distortions and illumination conditions not present in the training set.

The development of effective and editable material models is becoming increasingly important so that users of virtual prototyping, video-games or AR/VR applications can have compelling and realistic experiences. For instance, allowing for the creation of virtual environments that surpass the uncanny valley, or empowering artists to create breathtaking visual settings. Effective material representations require understanding the properties that uniquely define them, which we refer to as visual material attributes. These attributes are spatially-varying parameters that maintain spatial coherency with respect to the material structure while remaining invariant to changes in the scene illumination or the geometry of the underlying object. For example, they may represent optical properties of a microfacet spatially-varying BRDF (albedo, normals, roughness, anisotropy, etc.), but also artistic stylizations or higher-level properties, as in semantic segmentation masks. The former can be obtained with specific optical capture devices and estimation methods, either using machine learning [Mer+19] or classic optimization methods [Ste+14]. The latter may require time-consuming human labor.

Obtaining these attributes for large material samples can be problematic in certain situations. For instance, measuring high-resolution surface normals of large material samples or generating densely annotated segmentation masks, require either unfeasibly large capturing setups or expensive manual input. A common strategy to handle these cases consists of propagating these attributes using a larger image of the same material used as guidance image. This problem has been formulated within the context of image analogies [Her+01]. Propagating, or transferring, these properties requires preserving the local spatial regularities of the material, as well as adapting to the global variations of the guidance image, which is a highly challenging problem.

Previous work has used PatchMatch-based synthesis [Mel+12], look-up-tables [RPG16], or neural networks [Maz+19] by looking for repetitive patterns, match color statistics, or image gradients. These methods are however prohibitive due to their runtime performance, taking hours. Image analogies methods have been extended so as to leverage the power of convolutional neural networks as image descriptors [Lia+17; Ben+21], but those methods do not yield predictable outputs, and their computational complexity yields them impractical in interactive applications. Recently, neural networks have proven successful for the task of video stylization [Tex+20b], where the user needs to provide a few input editing exemplars which are propagated at interactive rates to the rest of the video. We show that this approach lacks the generalization power to work with images taken under unseen resolutions and geometric or illumination conditions.

In this chapter, we propose a novel learning-based method to propagate any visual material attribute śestimated locally for a materialś to larger samples of it. We train a neural network per material using image-to-image translation methods, making use of a policy of data augmentation that makes the transfer invariant to affine transformations (scale, rotations and shears). As opposed to other methods which blindly apply data augmentation strategies, our method is robustly evaluated using a real dataset to guarantee predictability in the estimations. Further, illumination invariance is obtained by feeding the network with multiple images of the material taken under a diverse set of illuminations. Our models can be trained in less than a minute and generalize to materials with the same microestructure.

In summary, we present the following contributions:

• The first method to use photometric data to train an image-to-image translation model capable of propagating any kind of visual material attribute to larger samples of the material, regardless of the illumination conditions of the input image.

• A data augmentation policy, thoroughly evaluated with a real dataset, designed to make the transfer invariant to affine deformations.

• Exhaustive comparisons with related work, demonstrating that we can achieve more predictable and higher-quality mappings with a fraction of the computational cost.

• We show that our trained models generalize to materials with similar microstructure as the ones used for training.

• On Section 2.C, we show an extension on this method for High Resolution SVBRDF capture, which is part of a publication accepted to SIGGRAPH 2023.

2.2 Related Work

Several computer graphics and vision problems are closely related to our method. The most similar ones are those related to any form of by-example visual attribute transfer (e.g color, texture, style, or geometry). Besides, we also review material estimation and capture methods.

2.2.1 Visual Attribute Transfer

Visual attribute transfer refers to the problem of transferring some visual attributes (e.g. color, style, texture, or geometry) of one or many exemplars to another exemplar while preserving its content.

This problem can be formulated within the context of Image Analogies [Her+01], in which the goal is to stylize a target un-stylized image B, given a pair of images A (un-stylized) and A’ (stylized). The most common approach to tackle this problem has been via patch-based texture synthesis [Bén+13; Bar+15; Jam+19]. More recent approaches have leveraged the capabilities of deep latent spaces within convolutional neural networks to disentangle style from content [Ree+15]. A seminal work by Gatys et al.[GEB15b] utilizes a VGG-19 convolu-tional neural network [SZ15] pre-trained on ImageNet [Den+09] as a feature descriptor for images, in which style and content are associated with different layers of the network, and transferred by gradient descent optimization. Their work on style transfer has been extended for single images [HB17; Li+17b; JAF16; Che+17b] and video [Che+17a], as well as for developing image-space distance metrics which resemble human perception [Zha+18]. A shared limitation of many of these methods is their narrow capabilities to provide predictable edits, and considerable focus nowadays is put towards this end [Gat+17; Gu+18]. A comprehensive review on the topic of neural style transfer is provided by Jing et al. [Jin+19]. Our work differs from traditional style transfer approaches in the sense that we deal with a more constrained problem that requires predictable outcomes.

Exploiting the power of deep neural networks in the image analogies problem was tackled by Liao et al. [Lia+17] who, by assuming a semantic prior over an exemplar input image and a target one, propose a method capable of finding a bijective mapping between both inputs, enabling two-way stylizations. Single-image generative models [SDM19] were extended to the image analogies problem in [Ben+21], by using convolutional neural networks to generate a new image with the style of an input style and the structure of another image. These methods, however, rely heavily on content or semantic features, making them vulnerable to lighting or geometric differences between the input images; and are computationally expensive, rendering them impractical for interactive applications. Our method is robust to both geometric distortions and illumination variations, and works at interactive rates. Similar in spirit to our method, as it explicitly considers texture variations due to illumination, is the work of Fišer et al. [Fiš+16], which applies patch-match to provide illumination-dependent exemplar-based stylizations to cartoon pictures. In contrast, our work is meant to be illumination-invariant. Also concerned with stylization problems, Texler et al. [Tex+20b] present a method highly related to ours. They apply patch-based training of an encoder-decoder deep neural network using key frames stylized by a user. Resembling our approach, their algorithm also follows a few-shot learning strategy using as training data a few exemplar patches. However, as opposed to our method, they do not account for the variability of the appearance of materials under different lighting and viewpoint so, as we show in Section 2.7, their method does not generalize to unseen illuminations or geometric variations.

Image colorization is concerned with colorizing a gray-scale image given a few colorized exemplars. In this problem, it is critical to infer semantic relationships between the images so that the new scene is perceptually coherent and plausible [He+18; Zha+19a; He+19]. Similarly, edit propagation methods [AP08; End+16] work by propagating strokes provided by a user to the rest of the image, removing the need for a semantic understanding of the input image. Our work is related to the latter techniques, as we leverage the feature spaces of the CNNs and do not require a large labeled dataset to effectively solve our problem, and it can also be used to propagate segmentation masks [LAA08].

2.2.2 Textured Materials

Many real-world materials show spatial regularities, commonly referred to as textures. The patterns present in textures can be parameterized, which allows for low-cost material capture or synthesis models. A way to model textured materials is through BTFs (Bidirectional Texture Functions) [Dan+99; LM01], a technique that uses multiple camera views and lighting angles to capture a dense sampling of the appearance of a material. Inspired by such methods, we leverage several images of the material under different illumination conditions, however, we require fewer data than typical BTFs capture setups [Rai+19; Rai+20]. Similarly, the problem of extrapolating BTFs captures to larger material samples was addressed by Steinhausen et al. [Ste+15a; Ste+15b], who propagate measured BTFs using texture synthesis. Our method is not meant to propagate full BTFs measurements but could potentially be applied to such datasets.

The goal of texture synthesis is to reconstruct a larger image given a small sample leveraging structural content. This is a long-standing problem in the computer graphics field and different strategies have been proposed, for instance, using PatchMatch [Dia+15], texture transport [AWL15], point processes [Gue+20; LH06], or neural networks [EM17; Zho+18; FAW19; Rod+19]. Also related to our work, Li et al. [LPG19] capture the appearance of materials by first estimating their BRDF and, then, synthesizing the high resolution microstructure from a dataset of measured SVBRDFs. Our problem is unlike texture synthesis, as we do not aim to create novel content but to predictably transfer visual material attributes.

2.2.3 SVBRDF Estimation

The problem of estimating a SVBRDF model from one or several images using lightweight capture setups is becoming increasingly popular in the literature. Early work [HS05; HS03] leveraged Photometric Stereo [Ike81] and SVBRDF manifold bootstrapping [Don+10] for surface geometry reconstruction, while newer methods exploit the power of deep neural networks. Recent surveys by Guarnera [Gua+16] and Dong [Don19] contain relevant approaches. While our method is not meant to estimate the SVBRDF properties of a material, it can be used in combination with those techniques to create larger material assets.

There are a few methods that follow a similar paradigm to ours, transferring pre-estimated SVBRDF maps to a larger material sample. Using PatchMatch texture synthesis, Melendez et al. [Mel+12] transfer displacement and albedo maps from small samples of the materials. Their method is limited to daylight illumination and materials present in façades. By means of look-up-tables, and using surface normals and specular as guidance, Riviere et al. [RPG16] transfer surface reflectance captured with controlled LCD lighting to a material sample observed under natural lighting. Recently, Deschaintre et al. [Des+19] fine-tune a network trained to estimate SVBRDFs [Des+18], to work on larger material samples taking a guidance image as input. This approach is limited to transfer a pre-defined set of property maps while our method can transfer any kind. The strategy of using multiple images of the material under different illuminations as input data is not new. Li et al. [Li+17a] and Ye et al. [Ye+18] utilize a self-augmentation strategy to make the estimation of the SVBRDF more robust to unknown environment illumination. We are inspired by these approaches to increase robustness in the model predictions.

2.3 Problem Formulation

Our goal is to transfer a D-dimensional spatially-varying visual attribute w of a material (for example, estimated locally at high resolution) to a larger sample of it.

We formulate the problem with an image-to-image translation approach. For training, our method takes as input: a photometric dataset IL, and a visual attribute w. The photometric dataset consisting of a number of RGB planar images of a material, IL={ililn×m×3}, |IL|1, illuminated with different light sources l ∈ L, of n × m pixels. This kind of images can be either captured with specific devices [Nam+16; MWK17; Alc+19], or synthetically rendered given an inverse material estimation pipeline [Guo+20b; DDB20]. The visual attribute being a spatially-varying map w ∈ ℝn×m×D of any kind, and dimensions D, that maintains pixel-wise correspondence with the photometric input images IL. Figure 2.3 and Figure 2.5 show examples of these images for three different visual attributes: a stylization, a segmentation, and a normal map.

Figure 2.3: Overview of the method. We learn a mapping between the photometric response IL of the material and a visual property map w. We make this mapping robust to affine transformations by means of a particular policy of data augmentation used for training. We learn one model M per material and visual attribute, which allows us to robustly evaluate the performance of the method under several transformations of the guidance images X. At evaluation time, M can have any size. The training of M per visual attribute w takes less than a minute.

Figure 2.3: Overview of the method. We learn a mapping between the photometric response IL of the material and a visual property map w. We make this mapping robust to affine transformations by means of a particular policy of data augmentation used for training. We learn one model M per material and visual attribute, which allows us to robustly evaluate the performance of the method under several transformations of the guidance images X. At evaluation time, M can have any size. The training of M per visual attribute w takes less than a minute.

Figure 2.4: The color augmentation policy takes advantage of the material structural regularities to make the transfer robust to different albedos. (a) A diffuse image of the portion of the material used for training, and its corresponding normal map. (b) Input guidance image. (c) Transferred normals without the color augmentation policy. (d) Transferred normals using color augmentation. Note that the training data does not include images containing the white yarn.

Figure 2.4: The color augmentation policy takes advantage of the material structural regularities to make the transfer robust to different albedos. (a) A diffuse image of the portion of the material used for training, and its corresponding normal map. (b) Input guidance image. (c) Transferred normals without the color augmentation policy. (d) Transferred normals using color augmentation. Note that the training data does not include images containing the white yarn.

Figure 2.5: An overview of our evaluation dataset. (a) Five example images of our high resolution captured data illuminated with diffuse light and four directional light sources. (b) Examples of some of our visual property maps: colorization, normals and yarn segmentation. (c) Evaluation images taken under diffuse and directional illumination sources. Denim, (c) contains labels for several patches used for the evaluation in Section 2.6. Those materials show different properties that may prove challenging for our visual attribute transfer task. Our denim and knit fabrics show diverse color variations (different colored dyed yarns and stochastic albedo, respectively), linen shows strong geometric variations and satin shows anisotropic optical behavior.

Figure 2.5: An overview of our evaluation dataset. (a) Five example images of our high resolution captured data illuminated with diffuse light and four directional light sources. (b) Examples of some of our visual property maps: colorization, normals and yarn segmentation. (c) Evaluation images taken under diffuse and directional illumination sources. Denim, (c) contains labels for several patches used for the evaluation in Section 2.6. Those materials show different properties that may prove challenging for our visual attribute transfer task. Our denim and knit fabrics show diverse color variations (different colored dyed yarns and stochastic albedo, respectively), linen shows strong geometric variations and satin shows anisotropic optical behavior.

Given this data for a single material, and a strategy of patch-based training, we learn a function M that can be applied to a new guidance image of the input material (or a similar one) XN×M×3 of any size N × M, to get its corresponding visual attribute ΩN×M×D:

M:XΩ,                  (2.1)

s.t.ilMω,ilIL                  (2.2)

At evaluation time, the guidance image X might contain different colors, scales, illumination, or affine distortions than the images used for training. Figure 2.3 shows an overview of the training and evaluation processes. In Section 2.6, we evaluate the conditions of the input guidance image upon which the method provides robust estimations.

2.4 Learning Framework

In this section, we describe the patch-based and data augmentation strategies used for training, the neural network design and loss functions, and the implementation details.

2.4.1 Patch-based Training

Each training step takes as input a pair of corresponding image patches taken from the photometric input IL, and the visual attribute w. Using this data alone already provides a good starting point for generalizing to unseen illumination setups. However, it is not sufficient in scenarios in which the guidance image contains variations due to image noise, a different scale, or any other affine distortion. In order to make the transfer invariant to these transformations, the network needs to be trained with the appropriate data. Data augmentation strategies are essential for reducing the amount of necessary data for training [SK19; Kar+20a], however, random strategies not taking into account the particular domain might degrade the quality of the prediction. We therefore follow a pre-defined data augmentation policy T (illustrated in Figure 2.3), where random operations are performed sequentially.

Color Augmentation Even if the microstructure of the material is homogeneous and can be measured using only a small patch of it, it may not be possible for the network to estimate its visual attributes for parts of the material which contain previously unseen colors. We correct this by randomly permuting the color channels of the photometric input (see Figure 2.4). As we show in Section 2.7, this data augmentation policy also helps make the model generalize the transfer to similar materials, by learning features that are more related to the structure of the material than to its color. Similar operations have been recently proposed for finding robust visual representations on self-supervised settings [Che+20b]. Only the photometric input is subject to this transformation.

Affine Transforms In order to allow for material editing applications in real images, the model should generalize to images taken under camera perspectives and geometry distortions different than those present in the fronto-planar images used for training. Many of such texture irregularities can be defined as an affine transform: translation, rotation, shear, or scaling. As CNNs are shift-invariant by design [Kau17], we propose to augment our datasets with random transformations for scalings, rotations, and shears. Those transforms can be efficiently performed on images through matrix multiplications. However, some visual property maps need to be treated specially, as the spatial transform might have a different behavior in 3D vector space, e.g. normals or tangents. In those maps, each pixel is a representation of a 3D vector. As such, we also perform the rotations and shear operations to these maps in 3D space, by multiplying each normal vector by the same affine transformation matrix applied to the 2D image. We perform the rotation around the Z axis, thus, assuming the camera sensor is parallel to the object plane.

Cropping Inspired by recent work on patch-based learning [Tex+20b; Par+20], during training, the network receives small patches of each input pair. Those patches are randomly cropped from the randomly augmented images, so the network receives a considerable amount of variations of the same material, thus making generalization possible.

2.4.2 Network Design

We follow a uni-modal image-to-image translation learning strategy, assuming there is only one correct mapping from input to output image. This approach is reasonable for the kind of transfers we test in this work. However, it might fail for ambiguous cases where there are multiple suitable outputs for the same input, for which multi-modal approaches [Zhu+17b] or GANs are more advisable, albeit harder to train. Specifically, our model M is a shallow U-Net network [RFB15] with 4 blocks of layers, containing a small number of trainable parameters, inspired on few-shot learning strategies [Wan+20d; Wan+19; Liu+19b]. This type of fully convolutional architecture is of common usage in image-space regression problems, due to its capability of efficiently learning patterns at different levels of abstraction thanks to its multi-scale design. Skip connections are added to enhance local details [RFB15; Iso+17]. The small number of parameters allows for faster training and inference, as well as reduced memory usage.

Note that, as opposed to previous work [DDB20] on material transfer, which is initialized from a pre-trained network, but inspired by single-sample image synthesis methods [Zho+18; SDM19], we train one network per material and visual property. This strategy, although increasing training time, is key for the following reasons: First, it allows us to better understand the generalization capabilities obtained through our data augmentation policies. Second, it guarantees predictability of the trained model as every feature learned by the network is specific for each material and visual property pairs. Finally, it removes potential problems of a biased dataset as there is no cross-material or cross-domain learning. If the input dataset does not contain enough variations of the material to represent the whole material, the model will fail in areas with unseen patterns. In those cases, having a pre-trained network may help as a material prior, as in [DDB20]. However, our method obtains comparable results to pre-trained methods, with a smaller computational footprint and with the additional flexibility of not needing an expensive dataset and large models. In practice, training a single network per material and attribute is not problematic as this process takes less than a minute.

Loss Function Choosing the appropriate loss function for a learning framework highly depends on the problem. For example, some methods [Iso+17] combine a per-pixel L1 metric with adversarial [Goo+14a] losses, as the latter allows for better semantic mappings in multi-modal learning scenarios, whilst the former allows for improved predictability. Texler et al. [Tex+20b] further includes a perceptual loss [Tex+20a], while using a render loss is common in methods that estimate material parameters from photos [Des+18; DDB20]. In our method, we show that using a L1 loss for training is enough for learning accurate and predictable mappings in our regression task, and binary cross-entropy loss for semantic segmentation, following standard practice in image segmentation [RFB15]. We found that the L2 loss function yields overly-smooth outputs, and perceptual loss functions like LPIPS [Zha+18] are more prone to artifacts than pixel-wise norm metrics. We include an ablation study of the impact of the loss function in the supplementary material of this chapter.

2.4.3 Implementation details

We use PyTorch [Pas+19] as the learning framework, Adam [KB15] for optimization, a learning rate of 0.002, and a batch size of 16. The training images are randomly augmented using uniform distributions by the following operations, in order: First, the photometric input is subject to the color augmentation policy. Then, both inputs and targets are randomly rotated by an angle in the [−90, 90] range, randomly sheared by an angle in the [−45, 45] range, and randomly rescaled in the [0.5, 2] range of scale factors. Then, patches of 128 × 128 pixels are randomly cropped during training to generate a large dataset of images.

All inputs are always standardized using their own mean and variance. Each M is trained for 1000 iterations, which takes around 1 minute on a single Nvidia 1080Ti GPU. Due to the fully convolutional nature of our models and their reduced number of trainable parameters, the guidance images X used for evaluation can be of arbitrary dimensions. We used images of up to 5000 × 5000 pixels for which evaluation takes around 150ms. We refer the reader to the supplementary material in this chapter for a comprehensive description and a diagram of this model, as well as further implementation details.

2.5 Dataset and Metrics

2.5.1 Dual-Resolution Captured Data

Our method is agnostic to the capture setup [Nam+16; MWK17; Alc+19], and it may work for any kind of input data (e.g. BTFs [LM01; Rai+19]) as long as the photometric images are pixel-wise aligned with each other and to the visual attribute. This dataset could also be created synthetically by rendering the outputs of any SVBRDF estimation method [DDB20]. In Section 2.7, we show results of our method using these acquisition pipelines.

However, for the purpose of this evaluation, we choose to work with real data. The main reason is that the data obtained with material capture devices poses extra challenges that are difficult to reproduce with render engines. For example, material irregularities, complex optical behavior, or distortions and color shifts introduced by the optical system that might cause the models to produce inaccurate estimations.

We create a dataset containing images of the same material taken with two imaging camera systems: a high-resolution camera that allows us to take pictures of 0.7 × 0.9 cm, with a resolution of 367 × 490 pixels and a macroscopic camera which provides images of 11 × 11 cm, with a resolution of 4800 × 4800 pixels. In terms of illumination, our setup has 27 different collimated light sources uniformly distributed across the hemisphere as well as diffuse illumination. We build a dataset of four different textile materials, whose complex optical behavior due to anisotropy, transmittance, directionality and microstructure [CLA19] turns them particularly challenging for synthesis and editing operations [Kam+16; Rai+19]. For training, we capture one image for each light source, making a total of 28 different images (Figure 2.5 (a)). For the evaluation set, we take one guidance image with diffuse illumination and another guidance image with a directional light source (Figure 2.5 (c)). The visual attributes (Figure 2.5 (b)) are generated automatically using photometric stereo [Ike81] in the case of normal maps and manually by artists in the cases of colorizations and segmentation masks.

2.5.2 Attribute-specific Metrics for Evaluation

For evaluating our models, we choose domain-specific distance metrics different from those they were trained with to better understand their generalization capabilities [TOB15]. For normal maps, we compute the cosine distance between ground truth and estimated maps, as it accounts for the geometric space in which normals lie. In the case of image segmentation, we evaluate the results using the Jaccard similarity coefficient (IoU) [Zho+19], which is well-suited for sparse segmentation tasks. Finally, we use the state-of-the-art metric Learned Perceptual Image Patch Similarity (LPIPS) to evaluate the quality of the colorizations, as it has been shown to outperform L2 norm for visual perception tasks [Zha+18].

2.6 Evaluation

We evaluate our method in different settings. First, we assess the type and amount of photometric input data necessary for the model to generalize to different illumination conditions and image distortions. Then, we test our data augmentation strategy for affine transformations.

2.6.1 Invariance to Input Illuminations and Distortions

In this set of experiments, we aim to evaluate if the model produces the same output after changing the training data and input guidance images. To this end, we measure both the impact of different illuminations and sizes of the photometric dataset, as well as variations of the input guidance images.

Photometric Input

In this experiment, we study the type of photometric input that makes the method invariant to different illumination conditions of the guidance image. We compare three models trained with different datasets: diffuseNet, which uses a single image illuminated with diffuse lighting; photometricNet, which takes as input 27 directional lights; and diphotoNet, which uses all the 28 sources. For evaluation, we use the macroscopic camera which captures a larger sample of the material at a lower resolution. We take two test guidance images with diffuse and directional lighting (Figure 2.5 (c)). In this experiment, all images are aligned, therefore, we follow a limited policy of data augmentation from Section 2.4.1, applying only color augmentation, rescaling, and crops, and leaving out rotations and shears for in-the-wild scenarios.

Figure 2.6 shows qualitative results of our method for a selection of patches for the denimmaterial. Figure 2.6 (a) shows that the best accuracy is in general obtained with diphotoNet, i.e. training with both directional and diffuse illuminations. This is reasonable, as the model is trained using the same illuminations used during test time. Conversely, diffuseNet, a network only trained using diffuse lighting is less capable of generalizing under any kind of illumination source. This suggests that following a photometric approach for training material synthesis models allows for better generalization capabilities. Finally, the results for photometricNet, which is trained only with directional lights shows similar accuracy as diphotoNet while proving generalization to unseen illuminations. For all the results shown in the chapter, we have chosen photometricNet as our model. This has an additional advantage from the usability perspective: it is relatively easy and cost-effective to build a capture setup (such as the flash light from a smartphone) or generate renders with directional lights, while it is considerably harder to recreate both types of illumination consistently.

Figure 2.6: Qualitative results of our models under different datasets and data augmentation configurations, for different inputs of the denim material. On the left (a), we show the results of our networks trained using different dataset configurations, under a guidance image taken with diffuse lighting (p0), and three crops of a guidance image illuminated with a directional light (p1, p2, p3) not present in the training set. Please refer to Figure 2.5 for the position of these crops on the larger guidance images. On the right (b), we show the results of our photometricNet, under different geometric distortions (rotations and shears) performed to its guidances images, taken under diffuse illumination.

Figure 2.6: Qualitative results of our models under different datasets and data augmentation configurations, for different inputs of the denim material. On the left (a), we show the results of our networks trained using different dataset configurations, under a guidance image taken with diffuse lighting (p0), and three crops of a guidance image illuminated with a directional light (p1, p2, p3) not present in the training set. Please refer to Figure 2.5 for the position of these crops on the larger guidance images. On the right (b), we show the results of our photometricNet, under different geometric distortions (rotations and shears) performed to its guidances images, taken under diffuse illumination.

Dataset Size Influence

Our goal in this experiment is to understand how many images [NJR15] taken under directional lights are needed in order to obtain the desired invariance to illumination.

We train different models using reduced versions of our full 27-image dataset, and compare their results to those of the photometricNet trained on the full directional dataset. More precisely, we train networks using 1, 3, 9 and 18 directional lights for each material and application in our dataset following the same reduced data augmentation policy described in the previous experiment. Instead of randomly selecting light sources around the hemisphere, we perform a more sensible light source sampling. Specifically, for each reduced dataset, there is at least one light that is as close as possible to the normal of the surface on which the material lies, thus giving more importance to frontal angles. We extend these selected lights to 3 for the reduced datasets with more than three lights. The rest of lighting sources are uniformly sampled around the hemisphere, up to a zenith angle of 70 degrees. The supplementary material contains a diagram of the distribution of lights in the hemisphere. Results are shown in Figure 2.7, where we see that adding more lights to the dataset monotonically increases the generalization of every model. However, adding lights consistently shows diminishing returns, which suggests that a capture setup with around nine lights might be enough for relatively accurate estimations. Similar findings are reported in multi-image SVBRDF estimation methods [Guo+20b; Des+19].

Figure 2.7: Error of the reduced photometricNets on our ground truth guidance image (taken with diffuse illumination, not present in the training dataset) for each material in the dataset and the three visual attributes and corresponding error metrics, which are close to the normal of the surface in which the material lies. The width of the lines indicates the standard variation across 5 repeated experiments, where different light directions were randomly chosen to form the training datasets. Please refer to the supplementary material for the position of each light source.

Figure 2.7: Error of the reduced photometricNets on our ground truth guidance image (taken with diffuse illumination, not present in the training dataset) for each material in the dataset and the three visual attributes and corresponding error metrics, which are close to the normal of the surface in which the material lies. The width of the lines indicates the standard variation across 5 repeated experiments, where different light directions were randomly chosen to form the training datasets. Please refer to the supplementary material for the position of each light source.

Image Degradations

As discussed in Section 2.5.1, real captured images may be subject to distortions, shifts and noise introduced by the optical capture system. To fully understand the robustness of our models with respect to these types of imperfections, we synthetically modify the saturation, contrast and noise present in the X guidance images. As shown in Figure 2.8, M is robust to these types of degradations, even when a fair amount of details are lost in X. These images were not part of the training dataset.

Figure 2.8: Output of our model under different distorted inputs: a change of saturation, contrast, and Gaussian noise (σ2 = 255). The number at the bottom is the cosine distance with respect to the original estimation. As we can see, the output is consistent with a very small error in all cases.

Figure 2.8: Output of our model under different distorted inputs: a change of saturation, contrast, and Gaussian noise (σ2 = 255). The number at the bottom is the cosine distance with respect to the original estimation. As we can see, the output is consistent with a very small error in all cases.

2.6.2 Equivariance to Affine Transforms

Our goal in this experiment is to analyze the equivariance of our model to affine distortions of the guidance images. A function f is said to be equivariant with respect to a transformation T if f(T(x))=T(f(x)). Inspired by [Ben+20], we measure the equivariance of our model M with respect to different affine transforms T performed to an input guidance image X, by computing their difference using the corresponding metric defined in Section 2.5.2 (d(M(T(X)),T(M(X))). To achieve this, we extend the augmentation policy in the previous experiments by adding random shears and rotations to the training process. We train the photometricNet with and without random shears and rotations for every visual attribute in our dataset and on the denim material. In the case of the normals, as described in Section 2.4.1, we also perform the shears and rotations in the geometric space in which normals lie.

We measure their robustness with respect to three different transformations T : rescalings, rotations and shears. As before, we use a guidance image taken under diffuse lighting as input for these experiments and attribute-specific distance metrics. Figure 2.6 (b) shows that not adding those transforms generates visual artifacts for rotations and shears. Interestingly, without affine augmentations, the models hallucinates vertical yarns, as it is the only type of data that it has seen as input. Figure 2.9 shows quantitative metrics for the range of transformations T in which we evaluated the models. Augmenting the training dataset with both shears and rotations generally provides the best results. Furthermore, this enhanced data augmentation policy improves the robustness of the models with respect to the scale of their inputs. Notably, applying random rotations during training appears to have a larger impact than random shears on the robustness of the models. This finding suggests that applying blind policies of data augmentation [ASE17; SK19; San+19; Che+20b] may not be an optimal strategy for some applications like ours, as the networks may not learn relevant features, or overfit to noise present in the training dataset. Finally, it is worth noting that every network has the same number of parameters, and is trained for the same number of iterations, which shows that the generalization capabilities of the models can be increased at no extra parameter cost.

Figure 2.9: Quantitative evaluation of the robustness of our networks with respect to different transformations T (rescaling, rotation and shearing) of networks trained under different data augmentation policies for the denim material. From top to bottom, we show results on: recoloring, segmentation, and normals estimation. Lower is better for each metric

Figure 2.9: Quantitative evaluation of the robustness of our networks with respect to different transformations T (rescaling, rotation and shearing) of networks trained under different data augmentation policies for the denim material. From top to bottom, we show results on: recoloring, segmentation, and normals estimation. Lower is better for each metric

2.7 Results and Comparisons

In this section, we first compare our method with related approaches on image stylization and large scale SVBRDF material transfer. Then, we show the capabilities of our method to generalize to similar materials to those in their training set and present its limitations.

2.7.1 Interactive Stylizations

In the first category, the method of Texler et al. [Tex+20a] allows artists to interactively edit a few keyframes of a video and propagate that edition to the rest of the video. As ours, they formulate this transfer problem by training an encoder-decoder network using patch-based learning. But, in contrast to our work, they do not perform any data augmentation policy outside of random cropping. For a fair comparison with such method, we compare two setups: 1) using a single image as training input, and 2) using the photometric dataset. Results are shown in Figures 2.2 and 2.10. As Texler’s method is not scale invariant, in both cases the training data provided for their model has the same scale as the images used for testing. Our models have been trained with the full policy of data augmentation. In the first setup (Figure 2.10), we use the diffuse illumination and hence compare their output with our diffuseNet output. As shown, none of the methods provide high quality results but our model manages to provide closer estimations. In the second setup (Figure 2.2), we train their model with the photometric input. The best results are obtained with our method. Even though extending the input data using photometric cues has an impact on the quality of Texler’s results, the lack of a data augmentation policy makes the transfer fuzzier and noisier. Further, their combination of style, adversarial and pixel-wise losses fails to yield predictable mappings. These results confirm the importance of a comprehensive data-augmentation policy, such as the one we propose, when using neural networks for image processing tasks of this kind. In addition, our model is trained in less time with a smaller computational footprint (1 minute vs 5 minutes).

Figure 2.10: Comparison of our method with Texler et al. [Tex+20b] using a single diffuse image of the denim material as training data. The task is to transfer the two attributes shown (Stylization and Segmentation) to two different guidance images. Even using a single image instead of a photometric dataset, we achieve higher quality mappings at a lower cost.

Figure 2.10: Comparison of our method with Texler et al. [Tex+20b] using a single diffuse image of the denim material as training data. The task is to transfer the two attributes shown (Stylization and Segmentation) to two different guidance images. Even using a single image instead of a photometric dataset, we achieve higher quality mappings at a lower cost.

Another way of formulating this visual attribute transfer problem is through image analogies. Using our single diffuse image for input, we compare our approach with two methods as shown in Figure 2.11. First, the work of Liao et al. [Lia+17], which uses deep latent spaces as image descriptors; and the method of Benaim et al. [Ben+21] that trains single-image generative models to find bijective mappings between the structure of one image and the style of another. Our method qualitatively outperforms these methods with a fraction of the computational cost: 1 minute in our case, 10 minutes in [Her+01], 40 minutes in [Lia+17] and 10 hours in [Ben+21]. Once trained, our models can be used to evaluate any guidance image in real time for materials with similar microstructure. In contrast, image analogies methods require expensive optimizations for each guidance image. We refer the reader to the supplementary material for more comparisons with these methods.

Figure 2.11: Comparison of image analogies approaches for input shown on the left, of the denim material. From left to right: Deep Image Analogies [Lia+17] and Structural Analogies [Ben+21]. Our results shows more accurate and predictable mappings, at less computational cost than the alternatives.

Figure 2.11: Comparison of image analogies approaches for input shown on the left, of the denim material. From left to right: Deep Image Analogies [Lia+17] and Structural Analogies [Ben+21]. Our results shows more accurate and predictable mappings, at less computational cost than the alternatives.

Figure 2.12 shows additional results of material stylizations. In these examples, we used the method of Gatys et al. [GEB15b] to stylize a small patch of the material. Then, we trained a model using photometricNet and diffuseNet. As the guidance image we used a bigger image with diffuse illumination. Compared with naïve style transfer applied to the whole image, our approach provides detailed stylizations where the microstructure of the material is preserved. We further see that diffuseNet provides noisier results than photometricNet, probably due to the fact that the photometric cues help to preserve the local shading variations.

Figure 2.12: Our method allows for interactive material-aware visual attribute transfers. Using an off-the-shelf style transfer algorithm [GEB15b], we can transfer the style of one image Istyle to the content of another Icontent, obtaining a visual attribute w. Training a Mω to learn this relationship, we can find predictable style mappings, that we can transfer to guidance images X, obtaining style transfers Mω(X). Learning this transfer is inexpensive and allows for interactive editions. Performing this transfer directly to the guidance image generates artifacts and not-predictable mappings.

Figure 2.12: Our method allows for interactive material-aware visual attribute transfers. Using an off-the-shelf style transfer algorithm [GEB15b], we can transfer the style of one image Istyle to the content of another Icontent, obtaining a visual attribute w. Training a Mω to learn this relationship, we can find predictable style mappings, that we can transfer to guidance images X, obtaining style transfers Mω(X). Learning this transfer is inexpensive and allows for interactive editions. Performing this transfer directly to the guidance image generates artifacts and not-predictable mappings.

2.7.2 Creation of Large Scale Digital Material Assets

Our method can be used to propagate SVBRDFs estimated locally in a small area of the material to larger samples. Figure 2.13 illustrates for a diverse set of materials that we can propagate albedo and normals estimated at high resolution in a small area of 0.7 x 0.9 cm to guidance images of 13 x 13 cm taken with a smartphone. Even though we train the albedo and normals models separately, which does not guarantee pixel-wise coherence between the estimated maps, the rendered images show realistic-looking materials, even with a diffuse material model. As opposed to the method of Deschaintre et al. [DDB20], which is trained to output directly Cook-Torrance [CT82] material layers, our method is agnostic to the parameters of the SVBRDF.

Figure 2.13: Results of our framework for material capture using a smartphone. Training two models, Ma and Mn with a photometric dataset, we obtain respectively albedo and normals from guidance images taken under uncontrolled conditions, which can be used by render engines. Here we have used Arnold [Geo+18] and a diffuse material model.

Figure 2.13: Results of our framework for material capture using a smartphone. Training two models, Ma and Mn with a photometric dataset, we obtain respectively albedo and normals from guidance images taken under uncontrolled conditions, which can be used by render engines. Here we have used Arnold [Geo+18] and a diffuse material model.

We assess the capabilities of our method on this setting using the same SVBRDF propagation scenario as proposed in [DDB20]. Using a small crop of a synthetic SVBRDF as input, we render 27 images using the same directional light position as we used in our real dataset, and train a photometricNet to estimate the surface normals from each of those renders. We then evaluate this model using a larger area of the material, illuminated under an unknown lighting position. In Table 2.1, we show a quantitative comparison with [DDB20], under different image quality metrics. As shown, our method achieves better scores on pixel-wise metrics, whilst [DDB20] achieves better deep perceptual scores, as in the LPIPS metric [Zha+18]. This might be related to the design of our loss function: we directly minimize pixel-wise differences, while [DDB20] is optimized using a render-aware loss. Qualitatively, as shown in Figure 2.14, our method obtains comparable quality mappings, with fewer artifacts. Despite capturing SVBRDF or arbitrary materials is not the goal of our method, we include in the supplementary comparisons with a direct SVBRDF acquisition method [Gao+19] using our own real dataset.

Table 2.1: Quantitative comparison with [DDB20], on the studied materials and different performance metrics. As shown, our method provides better pixel-wise accuracy than [DDB20], while their method obtains better perceptual scores.

Material ID

SSIM [Wan+04]↑

PSNR

MSE

LPIPS [Zha+ 18] ↓

[DDB20]

Ours

[DDB20]

Ours

[DDB20]

Ours

[DDB20]

Ours

560

0,905

0,925

31,190

32,310

0,002

0,002

0,244

0,221

1581

0,645

0,675

29,740

29,920

0,006

0,005

0,398

0,446

1684

0,654

0,746

29,711

30,410

0,007

0,003

0,196

0,251

2111

0,770

0,783

32,250

32,881

0,002

0,002

0,397

0,383

Average

0,744

0,782

30,723

31,380

0,004

0,003

0,309

0,325

Figure 2.14: Comparison of our method with the Guided Fine Tuning, by Deschaintre et al. [DDB20]. Following their algorithm, we render the Input SVBRDFs, and train a photometricNet on those renders. As shown, our method can achieve higher quality normal maps, with fewer artifacts. Input SVBRDFs, X, ground truth and results from Guided Fine Tuning were obtained directly from [DDB20].

Figure 2.14: Comparison of our method with the Guided Fine Tuning, by Deschaintre et al. [DDB20]. Following their algorithm, we render the Input SVBRDFs, and train a photometricNet on those renders. As shown, our method can achieve higher quality normal maps, with fewer artifacts. Input SVBRDFs, X, ground truth and results from Guided Fine Tuning were obtained directly from [DDB20].

A similar setting can be used to propagate attributes leveraging BTF measurements as training data. Figure 2.15 shows an example using a BTF from [WGK14]. Using Photometric Stereo [Ike81], we compute their surface normals and train a photometricNet on a central crop of the BTF, using the captures at a camera position of (ϕ = 0°, θ = 0°). We then evaluate this model using the full material surface, and a novel camera position, of (ϕ = 15°, θ = 11°). As shown, our model is capable of working with captured BTF data.

Figure 2.15: Our method can work with real BTF captured data. Training a photometricNet on a crop of the material (summarized in Training Dataset), to output surface normals, we can estimate surface normals on larger areas of the material, even under novel viewing positions.

Figure 2.15: Our method can work with real BTF captured data. Training a photometricNet on a crop of the material (summarized in Training Dataset), to output surface normals, we can estimate surface normals on larger areas of the material, even under novel viewing positions.

2.7.3 Generalization to Similar Materials

In previous experiments, we have shown the performance of our models when the guidance image corresponds to the material used for training. In this experiment, we show that our models, despite being trained only on a set of captures of a single material, generalize to materials of the same category. Figure 2.16 shows some examples for models trained using our dataset (Figure 2.5), taking input guidance images of different albedos and scales. The transfer works thanks to the network design and data augmentation strategy that is designed to use microstructure details as guidance. Our method could thus be used to transfer visual attributes for a diverse set of materials by simply training one model using a single but representative material of each category.

Figure 2.16: Generalization capabilities of our method when evaluated on materials similar to those in their training dataset. On the top row, we show the outputs of the model trained on the knit in our dataset (Figure 2.5), and evaluated on a different guidance images. We also show examples on linen and denim fabrics, with different conditions of saturation, blur, scale, and illumination. Even in very challenging cases, where the structure of the material is barely visible, as in the overly-saturated red linen or the noisy blue knit fabrics, our photometricNets can yield plausible results. The insets represent the training dataset by each model, as represented in Figure 2.5.

Figure 2.16: Generalization capabilities of our method when evaluated on materials similar to those in their training dataset. On the top row, we show the outputs of the model trained on the knit in our dataset (Figure 2.5), and evaluated on a different guidance images. We also show examples on linen and denim fabrics, with different conditions of saturation, blur, scale, and illumination. Even in very challenging cases, where the structure of the material is barely visible, as in the overly-saturated red linen or the noisy blue knit fabrics, our photometricNets can yield plausible results. The insets represent the training dataset by each model, as represented in Figure 2.5.

2.7.4 Limitations

Our models are not guaranteed to provide high-quality results outside the range of input data and data augmentation policies we train them on. This limitation is common to all learning-based approaches. It is unlikely that the framework is capable of generalizing to resolutions higher to those of the training data; or down-sampled images in which the texture details are not recognizable. Furthermore, the type of transformations we apply to the training data may not represent all the possible geometric variations, or non-linear warpings that materials are subject to in the real world. For materials which exhibit a strong variation in their microstructure, which cannot be fully captured using a single photometric dataset, our patch-based approach will likely fail to generalize to the full heterogeneity. Figure 2.17 shows two examples of failure cases for which the test data is not included in the training data and where the affine and illumination transformations are outside the suitable range.

Figure 2.17: Failure cases of our method. As shown in the first row, if the input dataset does not represent the heterogeneity present in the guidance X, the model fails to yield compelling results on unseen structures of the material, as shown on the green box. On the second row, we show a guidance image of the denim material X which exhibits strong geometric and illumination variations, outside of the range in which we train M with. As such, the model shows a poor performance on the segmentation task.

Figure 2.17: Failure cases of our method. As shown in the first row, if the input dataset does not represent the heterogeneity present in the guidance X, the model fails to yield compelling results on unseen structures of the material, as shown on the green box. On the second row, we show a guidance image of the denim material X which exhibits strong geometric and illumination variations, outside of the range in which we train M with. As such, the model shows a poor performance on the segmentation task.

2.8 Conclusions and Future Work

In this chapter, we have proposed a neural visual attribute transfer framework capable of transferring, for a given material, many types of visual property maps to images of unseen patches of the same -or similar- material taken under different illumination, capture setup, and affine distortion. To our knowledge, the proposed framework is the first method capable of leveraging the optical behavior of the material to this purpose by being trained using a photometric approach. Such an approach, besides the illumination-invariance we have shown, helps the neural network learn better mappings between visual domains, finding a physically-based representation of the material. Further, we have presented a comprehensive policy of data augmentation which outperforms previous work on visual attribute transfer given a single image of the material.

We have shown that our method can be used to transfer any kind of visual attribute estimated locally to larger material samples. Further, we have demonstrated that our models, although trained on a single material, generalize to materials of the same category. We think our findings will inspire future work showing that smart training strategies might alleviate the need for massive datasets. Our method could be extended in several ways. The need for obtaining high-resolution captures taken under different illuminations may be reduced by generating rendering images through recent advances in inverse material acquisition [Don19]. Further training with multiple patches may help to cope with material heterogeneity [Tex+20b]. Similarly, our findings suggest that extending the data augmentation policy to include 3D deformations will likely improve the accuracy. Beyond the generation of large scale digital assets for rendering, our method may have potential in other visual computing applications that require a low level understanding of the properties of the materials in real scenes. For example, the yarn segmentation application shown in the chapter might be suitable as input to shape from texture applications. Specific visual attributes might be useful to identify or highlight defects for image forensics problems, or to enhance different features in real-time or AR applications.

2.A Additional Implementation Details

Network design

Our model, summarized in Figure 2.18, is a standard U-Net [RFB15] encoder-decoder network with 4 blocks of layers and skip connections. U-Nets are widely used in pixel-wise regression and image-to-image translation problems, due to their efficient design, capable of exploiting texture and semantic patterns at different levels of abstraction, due to their multi-scale design. Skip connections are known to significantly enhance the quality of deep image-to-image translation models [Iso+17]. Each block is comprised of a tuple of convolutional layers with kernel sizes of 3 × 3, followed by a batch normalization operation [IS15] and a ReLU [NH10] non-linearity. The number of convolutional filters in each layer is shown in Figure 2.18. The last layer of the network uses 1 × 1 convolutional filters to transform the feature vectors into the desired number of classes (e.g 3 output channels for RGB regression tasks, as in normal map estimation or recoloring; or 1 output channel for the semantic segmentation problem). The total number of trainable parameters depends on the number of property maps that the network should jointly learn, but are around 483000 for all the problems we showcased in this work. Trainable weights are initialized randomly, by sampling a N(μ=0,σ=0.02). The output of each block in the encoder part of the network is max-pooled to halve its spatial resolution. Upsampling in the decoder side of the network is done through up-convolutions. For image segmentation problems, we add a Sigmoid non-linearity to the output of the network. Due to the reduced number of trainable parameters, models take only 1.94Mbs, allowing for fast training and evaluation. On Table 2.2, we show a quantitative study on the influence of the depth of the model in its accuracy, using the 1684 material from [Des+18] as the training dataset. A visualization of these results is included in Figure 2.19, showing that deeper models yield overly-smooth results for this few-shot learning problem.

Figure 2.18: An overview of our deep learning architecture Mω, which is a shallow U-net net-work [RFB15]. Below each block of layers, we show the number of trainable convolutional kernels, as well as the spatial resolution spanned by each block, with respect to the input resolution D. We use the normal map estimation problem for a denim fabric for visualization purposes.

Figure 2.18: An overview of our deep learning architecture Mω, which is a shallow U-net net-work [RFB15]. Below each block of layers, we show the number of trainable convolutional kernels, as well as the spatial resolution spanned by each block, with respect to the input resolution D. We use the normal map estimation problem for a denim fabric for visualization purposes.

Table 2.2: Quantitative comparison between different number of blocks of layers in our model. Every model we use throughout this work has 4 blocks of layers, as it provides the best trade-off between computational cost and precision. Deeper models fail to generalize properly in this few-shot learning scenario. We use a color code to highlight best and worst cases

Depth

SSIM [Wan+04] ↑

PSNR ↑

MSE ↓

LPIPS [Zha+18] ↓

Size (Mbs) ↓

Train Time (s) ↓

3

0.71

29.83

0.0052

0.37

0.49

51

4

0.74

30.41

0.0031

0.25

1.94

58

5

0.72

30.37

0.0043

0.33

7.63

164

6

0.72

30.87

0.0041

0.31

30.49

283

Figure 2.19: Impact of the depth of the model on its accuracy, using a synthetic dataset. As shown, our baseline outperforms deeper and shallower configurations.

Figure 2.19: Impact of the depth of the model on its accuracy, using a synthetic dataset. As shown, our baseline outperforms deeper and shallower configurations.

Training details

We use PyTorch [Pas+19] as the learning framework, Adam [KB15] for optimization, a learning rate of α=0.002,β1=0.9,β2=0.999,ϵ=10e8, and a batch size of 16. The training images are randomly augmented using a uniform distribution by the following operations, in order: randomly rotated by an angle in the [−90, 90] range, randomly sheared by an angle in the [−45, 45] range, and randomly rescaled in the [0.5, 2] range of scale factors. Their color channels are randomly reordered and patches of 128 × 128 pixels are randomly cropped during training, to generate a large dataset of images. The rescaling factors and rotation angles are randomly chosen for each image in each batch using a uniform distribution on the mentioned ranges. All these data augmentation operations are also applied to the target visual attributes, except for the color channel reordering, which is only applied to the input images. Images and visual attributes are always standardized using their own mean and variance. All the data augmentation operations are performed natively in GPU, which allows for faster training. To further reduce the computational cost and training times of our models, we train our models using automatic mixed precision training [Mic+18], which reduces the memory footprint of the models in GPU, while allowing for faster computation of certain operations, such as convolutions.

Each Mω is trained independently for each material and application (no cross-domain training or fine-tuning) for 1000 iterations, which takes around 1 minute on a single 1080Ti GPU. Due to the fully convolutional nature of our models, and their reduced number of trainable parameters, the guidance images X used for evaluation can be of arbitrary resolutions. We evaluate the models using half-precision, which allows us to utilize guidance images of up to 5000 × 5000 pixels, for which evaluation takes around 150ms in GPU. Thanks to the reduced number of parameters, an efficient use of GPU-native data augmentation techniques and leveraging half precision training and evaluation, we are able to train and evaluate our photometricNets in less than a minute of total computational time.

Ablation Study of the Loss Function

Finally, in Figure 2.20, we show a comparison of the results of our models when trained on different loss functions. As shown, even though the differences are sometimes subtle, the L1 consistently yields sharper estimations than the L2, whilst the perceptual loss function is more artifact-prone and thus less predictable. L1 is a common loss function in many image regression tasks, as in image-to-image translation [Iso+17] or in image synthesis [Zho+18].

Figure 2.20: Impact of the choice of the loss function used to train the networks in the estimated maps at macroscale. On the left, input guidance X and visual attribute w to be transferred. On their right, the result of Mω when trained using L1, L2 and LPIPS [Zha+18] loss functions. All the networks in this experiment were trained with the same configuration, only the loss function was modified.

Figure 2.20: Impact of the choice of the loss function used to train the networks in the estimated maps at macroscale. On the left, input guidance X and visual attribute w to be transferred. On their right, the result of Mω when trained using L1, L2 and LPIPS [Zha+18] loss functions. All the networks in this experiment were trained with the same configuration, only the loss function was modified.

Additional Dataset Details

We show the positions of the lights in the photometric dataset in Figure 2.21.

Figure 2.21: Relative position of the 27 directional light sources used to capture the photometric datasets with respect to the material sample, in degrees. In blue, we show the three lights that are always included in our reduced photometric datasets for the light sampling experiments. As shown, the lights cover a wide variety of azimuth and zenith angles, allowing for an accurate capture of the optical behavior of the materials.

Figure 2.21: Relative position of the 27 directional light sources used to capture the photometric datasets with respect to the material sample, in degrees. In blue, we show the three lights that are always included in our reduced photometric datasets for the light sampling experiments. As shown, the lights cover a wide variety of azimuth and zenith angles, allowing for an accurate capture of the optical behavior of the materials.

2.B Additional Comparisons

Comparison with Image Analogies Methods

In Figure 2.22, we compare our method to different image analogies approaches. We show a comparison between [Her+01], [Lia+17], [Ben+21], and our method. The seminal work in image analogies [Her+01] fails in this visual attribute transfer problem, most likely due to the difference in scale between input and target images. For the comparisons with [Lia+17] and [Ben+21], we reduced the scale of the target images D to 256 × 256 pixels to make their methods computationally tractable. The methods in [Lia+17; Ben+21] never see the diffuse image A, as they generate an intermediate representation by themselves. While showing better results than [Her+01], methods which use deep learning for this task still fail to yield compelling or predictable results. The method in [Ben+21], which leverages structural information in the images to train a single-image generative model [SDM19], shows more plausible results, but with an expensive computational footprint (10 hours per network), compared to the 45 minutes in [Lia+17], 10 minutes in [Her+01] and less than 1 minute in our case. Because we assume pixel-wise semantic correspondence between input image and visual attributes, we can find mappings in interactive times, and which are significantly more predictable and of higher quality than these image analogies methods.

Figure 2.22: Comparison of photometricNet with related image analogies methods. In order, we show a comparison between [Her+01], [Lia+17], [Ben+21], and our method. Our method clearly outperforms any analogies method, at a fraction of the computational cost. The comparisons with [Her+01] were done using an unofficial implementation of their method.

Figure 2.22: Comparison of photometricNet with related image analogies methods. In order, we show a comparison between [Her+01], [Lia+17], [Ben+21], and our method. Our method clearly outperforms any analogies method, at a fraction of the computational cost. The comparisons with [Her+01] were done using an unofficial implementation of their method.

Comparison with Texler et al.

In Figure 2.23, we show a comparison with different configurations of [Tex+20b]. Specifically, we train their method using a dataset with a single image using diffuse illumination (left) and the full photometric dataset (middle column). Their method is trained individually for each dataset and visual attribute, for around 5 minutes to guarantee convergence, as suggested in their paper. As [Tex+20b] assumes that the scale (pixels/cm) of the dataset and evaluation images are the same, we rescale our input dataset to match the scale of the evaluation images. We evaluate each method using three different images X, one with diffuse-light illumination (top row), another with diffuse-light illumination but with a geometric transformation (middle row), and an image captured with directional lighting (bottom row). As shown, by enhancing the method in [Tex+20b] with a photometric dataset, their model is able to generalize better to every image X in the evaluation dataset than its counterpart trained only using a single image. However, our method, specifically designed for this illumination and geometry invariant visual attribute transfer outperforms their work in every type of input and visual attribute, whilst maintaining a lower computational overhead (1 minute of training in our case) and memory footprint (their trained network uses up 12 Mbs, while ours is 1.9 Mbs). Our model trained using a single illumination (diffuseNet) shows better results than their method trained using the photometric dataset, suggesting that our transfer method has more adequate inductive biases for the visual attribute transfer problem.

Figure 2.23: Comparison of our method with different configurations of [Tex+20b]. On the top row, we compare their results and ours using a dataset containing only one image, captured with diffuse lighting. On the bottom row, we show results when training these methods with a full photometric dataset. Our method is able to generalize better to different illumination and geometric conditions, with a smaller computational overhead.

Figure 2.23: Comparison of our method with different configurations of [Tex+20b]. On the top row, we compare their results and ours using a dataset containing only one image, captured with diffuse lighting. On the bottom row, we show results when training these methods with a full photometric dataset. Our method is able to generalize better to different illumination and geometric conditions, with a smaller computational overhead.

Comparison with deep SVBRDF acquisition

In Figure 2.24, we show a comparison with a recent deep SVBRDF acquisition method [Gao+19], which relies on a pre-trained autoencoder and a differentiable render engine for optimizing the initial estimation of the autoencoder. As shown, our method generates sharper and more coherent normal maps. PhotometricNet relies on a small sample of the same material for training, and is thus less sensitive to potential biases of the training dataset, with the cost of having reduced generalization capabilities compared to multi-material methods.

Figure 2.24: Comparison of our method with [Gao+19]. From left to right, we show the input guidance X, the output of photometricNet Mω trained to transfer normal maps, and the output of [Gao+19] after 5000 optimization steps. Their method relies on a pre-trained network, which learns a prior over materials that does not generalize to examples outside their training set. The images in this experiment are of 256 × 256 pixels, to match the input requirements in [Gao+19].

Figure 2.24: Comparison of our method with [Gao+19]. From left to right, we show the input guidance X, the output of photometricNet Mω trained to transfer normal maps, and the output of [Gao+19] after 5000 optimization steps. Their method relies on a pre-trained network, which learns a prior over materials that does not generalize to examples outside their training set. The images in this experiment are of 256 × 256 pixels, to match the input requirements in [Gao+19].

2.C Extension to SVBSDF Propagation

In this section, we present an extension to the previous method for propagating material maps. With it, we can propagate multiple property maps at the same time, learn from many photometric datasets simultaneously, and achieve more accurate and sharper results. With these improvements, we are able to propagate multiple SVBSDF maps at the same time, allowing for efficient high-resolution material generation, as we show in Figures 2.25 and 2.26. These findings are part of the data generation process which is used to build the datasets we use to train the models described in Section 5 and [Rod+23a], and are also part of a paper on high resolution material digitization, which has been accepted to SIGGRAPH 2023:

Figure 2.25: Renders of SVBSDF maps transferred by our extension of photometricNet.

Figure 2.25: Renders of SVBSDF maps transferred by our extension of photometricNet.

Figure 2.26: Left, virtual scene rendered using eight different materials digitized with our optical device. Right, real photos of four of the materials taken under diverse illumination conditions: area lights, diffuse lighting, and directional lighting of a high-resolution patch.

Figure 2.26: Left, virtual scene rendered using eight different materials digitized with our optical device. Right, real photos of four of the materials taken under diverse illumination conditions: area lights, diffuse lighting, and directional lighting of a high-resolution patch.

“Towards Material Digitization with a Dual-scale Optical System” Elena Garces, Victor Arellano, Carlos Rodriguez-Pardo, David Pascual-Hernandez, Sergio Suja, Jorge Lopez-Moreno

ACM Transactions on Graphics (Proceedings of SIGGRAPH) (2023)

2.C.1 Mesoscale Propagation

To propagate material properties, a common solution is to utilize the estimation of material properties in a localized area and extended it to larger samples [ND06; RSK13; Ste+14; DDB20]. We follow this approach to propagate SVBSDF measurements, and build upon the work presented in Section 2 and [RG21], which leverage images of the material taken under different illumination conditions to train a neural network that propagates individual maps.

As input to this propagation step, we use a photometric dataset obtained with a microscopic setup, spatially-varying parameters of a micro SVBSDF (Illustrated in Figure 2.28), and a guidance image that will serve as a reference for the mesoscale propagation. We use as guidance the image taken with a mid-range camera and diffuse lighting. Unlike the previous approach, aimed at transferring individual maps, we introduce some modifications to jointly propagate the full stack of parameters. We further adapt the method to be able to take as input more than one micro SVBSDF, necessary if the variability of material is not visible in a single micro capture (see Figure 2.27 (a)). We propose the following modifications:

1. Larger Training Dataset. We improve the generalization of the model by using more training data in two ways: First, we use a larger photometric dataset, using more directional-lit images than what is described in 2.21, as well as diffuse lighting. Second, we use random Gaussian blurs for data augmentation to improve robustness to potential degradation in mesoscale guidance images, and remove affine distortions to the data augmentation policy as they are not required for our SVBSDF transfer problem. Additionally, we train the model for more iterations, with a larger batch size.

2. Enable Large Spatially-varying Materials. We account for materials in which the appearance variability is not observed in a single micro capture by training the propagation network with as many micro SVBSDF as necessary. Figure 2.27 (a) showcases an example that required two microscale captures. Further, we disable the color-invariance data augmentation that randomly shuffles color channels when multiple micro SVBSDFs are available.

3. Improved Architecture and Losses. Finally, we observed that the original architecture was not able to effectively transfer a large number of parameters at the same time. Therefore, inspired by recent work on intrinsic decomposition [Jan+17], texture synthesis [RG22], and material estimation [DLG21], we use a U-Net [RFB15] with a single encoder but a separate decoder for every map we aim to transfer. Further, we introduce residual connections [He+16; Dia+20] into every layer, and substitute Batch Normalization [IS15] for Group Normalization [WH18]. When multiple microscale captures are available, we use a model with additional filters in every layer, to better learn from these extended datasets. This design generates sharper texture maps which preserve the statistics and appearance of each fitted microscale map more accurately. We also added a multi-channel perceptual component [CHB21] to the loss function, which is a powerful regularizer for texture synthesis. Figure 2.27 (b-c) shows the difference with respect to the previous approach.

Figure 2.27: Comparison with photometricNet [RG21] for mesoscale maps propagation. The first row (a) showcases an example where multiple captures at microscale were needed to cover the spatially-varying albedo of the material. (b) and (c) required a single capture, we show the result of normals and tangents maps compared with previous work.

Figure 2.27: Comparison with photometricNet [RG21] for mesoscale maps propagation. The first row (a) showcases an example where multiple captures at microscale were needed to cover the spatially-varying albedo of the material. (b) and (c) required a single capture, we show the result of normals and tangents maps compared with previous work.

Figure 2.28: On the left, microscale SVBSDF maps. On the right, a render generated with the maps propagated with our method.

Figure 2.28: On the left, microscale SVBSDF maps. On the right, a render generated with the maps propagated with our method.

2.C.2 Implementation Details

Problem Formulation

As in [RG21], our goal is to transfer spatially-varying attributes of a material captured at microscale to larger samples of the same material. We thus formulate this problem using an image-to-image translation framework. For training, our methods take as input a set of D photometric datasets of a material ILD, and their pixel-wise corresponding SVBSDFs MPD, each comprised of a list of p maps: p ∈ {albedo, roughness, transmittance, IOR, anisotropy, tangents, normals, specularTint, opacity}. Each photometric dataset is comprised of a set of RGB images of the material captured under different illumination conditions: ILd={ildild Rn×m×3},|ILd|1,dD. In contrast to [RG21], which only allowed to learn from a single photometric dataset and to transfer a single attribute, our extended approach allows for learning from multiple datasets |D| ≥ 1 and multiple property maps |p | ≥ 1 using a single model. We train a model T which learns to transfer from each image in each photometric dataset to its corresponding SVBSDF: T(ild)MPd,dD,lL. During inference, we use a l p guidance image of the material X, which represents a larger sample of it, and estimate its corresponding property maps: MXT(X). At evaluation time, this guidance image X may be captured with different conditions to those in the training dataset, including different camera or illumination conditions.

Datasets

We use a more comprehensive photometric dataset to what is proposed in [RG21]. In particular, for a single microscale capture, we use all the light sources used in the fitting algorithm, as well as diffuse-lit microscale images. Additionally, we use both polarization modes separately and, for the same light source, we construct an extra image by averaging the images taken under both polarization configurations. The total size of each photometric dataset is of |ILd|=108,dD, four times more data compared with the more limited 27 images proposed in [RG21]. For evaluation, our guidance image is taken with a mid-range camera and diffuse illumination, at a resolution of 4072 × 4072 pixels and a surface area of 10 × 10 centimeters. We typically use a single photometric dataset (|D|=1). However, for multi-colored or highly heterogeneous materials in which a single capture is not enough to fully represent the material variability, we capture as many datasets as needed.

Loss Function

Our loss function is the combination of two terms: weighted pixel-wise losses for each target map, and a multi-map style loss:

L=pλpLpixelp+λstyleLstyle                  (2.3)

Lpixel is the L1 norm weighted per map in the SVBSDF, λp. L1 produces sharper results than higher-order alternatives [RG21], such as L2. For the Lstyle loss term, we follow recent work on multi-channel texture synthesis, which extends deep perceptual losses for style transfer to the SVBRDF synthesis problem [CHB21]. This component of the loss function acts as a regularizer and allows the model to generate maps that better preserve the appearance of the ground truth training data.

Training Implementation Details

Model Design We use a lightweight U-Net [RFB15] architecture, with a few modifications based on recent work to maximize its efficiency and the quality of its outputs. We specify the full model architecture and model sizes in Figure 2.29. In every convolutional block of the model, we use residual connections [He+16; Dia+20], for better training convergence and preservation of details in the input images. We use 1 × 1 convolutions on these residual connections. We use a single decoder for each map, so as to maximally preserve their individual appearance and statistics. This has been proposed in recent work on texture synthesis [RG22], material capture [DLG21] or intrinsic decomposition [Jan+17]. On the convolutional blocks, we leverage Group Normalization [WH18]. Upsampling is done using transposed convolutions. As shown in Figure 2.29, the model has four hidden layers on the encoder and each decoder, however, we vary the size of those layers depending on whether the model is trained with a single microscale capture (for which we use a width factor of W = 16) or multiple (W = 32). Every other implementation detail in the model (stride, bias, pooling) follows [RFB15; RG21].

Figure 2.29: A full diagram of the model architecture we use for our mesoscale propagation problem. Building upon [RG21], use a lightweight U-Net [RFB15] architecture. Following previous work [Jan+17; RG22; DLG21], we use a different decoder for every map we aim to transfer. Each of our convolutional blocks, illustrated on the right, contain residual connections [He+16; Dia+20] and Group Normalization [WH18]. In red, we show the input/output dimensions (spatial, channels) of each layer; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, regularizations, and non-linearities. We set W to W = 16 when a single microscale capture is available as training data, W = 32 otherwise.

Figure 2.29: A full diagram of the model architecture we use for our mesoscale propagation problem. Building upon [RG21], use a lightweight U-Net [RFB15] architecture. Following previous work [Jan+17; RG22; DLG21], we use a different decoder for every map we aim to transfer. Each of our convolutional blocks, illustrated on the right, contain residual connections [He+16; Dia+20] and Group Normalization [WH18]. In red, we show the input/output dimensions (spatial, channels) of each layer; in blue, convolutional blocks and layers; in green, upsampling and concatenating operations; in yellow, normalization layers; and in purple, regularizations, and non-linearities. We set W to W = 16 when a single microscale capture is available as training data, W = 32 otherwise.

Training and Evaluation We use PyTorch [Pas+19], Torchvision [MR10], and Kornia [Rib+20] for training. We leverage mixed precision training and automatic gradient scaling [Mic+18], to accelerate the training process and regularize the models. Optimization is done using Adam [KB15] for 5000 iterations, with a learning rate of 0.002, a batch size of 40, and no weight decay. This process takes around 15 minutes when training with a single microscale capture, and 25 when using multiple captures as input. We evaluate the guidances images using half precision. We measure these times on an NVIDIA 1080Ti GPU.

Data Augmentation We follow the same training procedure defined in [RG21]. However, we do not use random rotations or shears, and remove the color invariance data augmentation whenever multiple microscale captures are available. Further, we introduce random Gaussian Blurs for data augmentation, using a p = 0.5, a kernel size of 5, and sigma selected uniformly at random for each element in each batch: σU(0.1,11). As in [RG21], we also perform random cropping, using crops of 128 × 128 pixels. For each element in each batch during training, we randomly choose the photometric dataset d ∈ D, and, from it, a random light source lL.

Loss Function We weight the loss function as follows: For the pixel wise loss: λnormals= 3,λIOR=1,λroughness=1,λalbedo=3,λtangent=3,λanisotropy=1,λspecularTint=1,λtransmittance=3,λopacity=1. The perceptual component [CHB21] is weighted with λstyle =0.25, and we use the AlexNet variant of LPIPS [Zha+18] as our backbone, which provides a powerful yet lightweight loss for texture transfer, as shown in [RG21].

2.C.3 Results

In Figure 2.30, we show an ablation study on the improvements of the maps propagation method proposed in [RG21]. We build upon their proposed implementation, removing their random shifts and rotations, and make progressive changes, in order, to the dataset size, model architecture, loss function, and data augmentation. We observe high-quality propagations, but the baseline sometimes produces overly smooth outputs (Fleece, Leather-Brown, Rib-Silver, or Leather-Lizard). Our proposed increased dataset and model architecture modifications enhance the maps’ sharpness, and our perceptual loss function allows for the propagation of additional details. Our final model removes the color augmentation proposed in [RG21] whenever multiple microscale captures are available, and introduces random blurs during training, allowing for achieving the highest quality results and for accurate base color propagations in challenging cases, as in Tartan.

Figure 2.30: Ablation study of our proposed improvements with respect to tphotometricNet [RG21], on base color (top rows), normals (mid) and tangent map (bottom) transfer. We start from the implementation in [RG21], which shows adequate but smooth results. On its right, we show the impact of increasing the dataset size, which tends to generate higher-quality estimations (Leather-Brown). With our improved architecture (fifth column), we achieve sharper maps (see Leather-Lizard, Rib-Silver, or Fleece). Using a perceptual loss (sixth column), we further improve the maps quality (Leather-Brown). Our final model (last column), which removes the color augmentation in [RG21] and introduces random blurs to the model training, achieves the highest quality, and allows for accurate base color transfers, as shown in Tartan. Input guidance images and training micro maps are shown on the first and second columns, respectively.

Figure 2.30: Ablation study of our proposed improvements with respect to tphotometricNet [RG21], on base color (top rows), normals (mid) and tangent map (bottom) transfer. We start from the implementation in [RG21], which shows adequate but smooth results. On its right, we show the impact of increasing the dataset size, which tends to generate higher-quality estimations (Leather-Brown). With our improved architecture (fifth column), we achieve sharper maps (see Leather-Lizard, Rib-Silver, or Fleece). Using a perceptual loss (sixth column), we further improve the maps quality (Leather-Brown). Our final model (last column), which removes the color augmentation in [RG21] and introduces random blurs to the model training, achieves the highest quality, and allows for accurate base color transfers, as shown in Tartan. Input guidance images and training micro maps are shown on the first and second columns, respectively.

BIBLIOGRAPHY

[Aba+16] Martın Abadi, Paul Barham, Jianmin Chen, Zhifeng Chen, Andy Davis, Jeffrey Dean, Matthieu Devin, Sanjay Ghemawat, Geoffrey Irving, Michael Isard, et al. “Tensorflow: A System for Large-Scale Machine Learning”. In: 12th Symposium on Operating Systems Design and Implementation. 2016, pp. 265–283 (cit. on p. 96).

[Aga+03] Sameer Agarwal, Ravi Ramamoorthi, Serge Belongie, and Henrik Wann Jensen. “Structured Importance Sampling of Environment Maps”. In: ACM Transactions on Graphics (TOG) 22.3 (2003), pp. 605–612 (cit. on p. 185).

[AAL16] Miika Aittala, Timo Aila, and Jaakko Lehtinen. “Reflectance Modeling by Neural Texture Synthesis”. In: ACM Transactions on Graphics (TOG) 35.4 (2016), pp. 1–13 (cit. on p. 124).

[AWL13] Miika Aittala, Tim Weyrich, and Jaakko Lehtinen. “Practical SVBRDF Capture in the Frequency Domain”. In: ACM Transactions on Graphics (TOG) 32.4 (2013), pp. 110–1 (cit. on pp. 17, 107, 128, 258).

[AWL15] Miika Aittala, Tim Weyrich, and Jaakko Lehtinen. “Two-shot SVBRDF Capture for Stationary Materials”. In: ACM Transactions on Graphics (TOG) 34.4 (2015), pp. 1–13 (cit. on pp. 17, 34, 104, 125, 258).

[Akl+18] Adib Akl, Charles Yaacoub, Marc Donias, Jean-Pierre Da Costa, and Christian Germain. “A Survey of Exemplar-Based Texture Synthesis Methods”. In: Computer Vision and Image Understanding 172 (2018), pp. 12–24 (cit. on p. 73).

[Alc+19] Raul Alcain, Carlos Heras, Iñigo Salinas, Jorge Lopez-Moreno, and Carlos Aliaga. “Microscale Optical Capture System for Digital Fabric Recreation”. In: Proceedings of the 7th International Conference on Photonics, Optics and Laser Technology - Volume 1: PHOTOPTICS, INSTICC. SciTePress, 2019, pp. 114–119 (cit. on pp. 35, 39).

[Alm+18] Amjad Almahairi, Sai Rajeshwar, Alessandro Sordoni, Philip Bachman, and Aaron Courville. “Augmented Cyclegan: Learning Many-to-Many Mappings from Unpaired Data”. In: International Conference on Machine Learning (ICML). PMLR. 2018, pp. 195–204 (cit. on p. 85).

[Ami+20] Alexander Amini, Wilko Schwarting, Ava Soleimany, and Daniela Rus. “Deep Evidential Regression”. In: Advances in Neural Information Processing Systems 33 (2020), pp. 14927–14937 (cit. on p. 126).

[AP08] Xiaobo An and Fabio Pellacini. “AppProp: All-pairs Appearance-Space Edit Propagation”. In: ACM Transactions on Graphics (TOG) 27.3 (2008), pp. 1–9 (cit. on p. 33).

[And+20] Pontus Andersson, Jim Nilsson, Tomas Akenine-Möller, Magnus Oskarsson, Kalle Åström, and Mark D Fairchild. “FLIP: A Difference Evaluator for Alternating Images.” In: Proc. ACM Comput. Graph. Interact. Tech. 3.2 (2020), pp. 15–1 (cit. on pp. 198, 199, 202).

[ASE17] Antreas Antoniou, Amos Storkey, and Harrison Edwards. “Data Augmentation Generative Adversarial Networks”. In: arXiv preprint arXiv:1711.04340 (2017) (cit. on p. 45).

[ACB19] Devansh Arpit, Vıctor Campos, and Yoshua Bengio. “How to Initialize Your Network? Robust Initialization for Weightnorm & Resnets”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on pp. 109, 119).

[ARV19] Yuki M Asano, Christian Rupprecht, and Andrea Vedaldi. “A Critical Analysis of Self-Supervision, or What We Can Learn from a Single Image”. In: International Conference on Learning Representations (ICLR). 2019 (cit. on p. 75).

[Att+22] Benjamin Attal, Jia-Bin Huang, Michael Zollhöfer, Johannes Kopf, and Changil Kim. “Learning Neural Light Fields with Ray-space Embedding”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 19819–19829 (cit. on pp. 187, 203, 212).

[Azi+19] Dejan Azinovic, Tzu-Mao Li, Anton Kaplanyan, and Matthias Nießner. “Inverse Path Tracing for Joint Material and Lighting Estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 2447–2456 (cit. on p. 203).

[Azi+23] Shekoofeh Azizi, Simon Kornblith, Chitwan Saharia, Mohammad Norouzi, and David J Fleet. “Synthetic Data from Diffusion Models Improves ImageNet Classification”. In: arXiv preprint arXiv:2304.08466 (2023) (cit. on p. 212).

[BKH16] Jimmy Lei Ba, Jamie Ryan Kiros, and Geoffrey E Hinton. “Layer Normalization”. In: arXiv preprint arXiv:1607.06450 (2016) (cit. on pp. 109, 111, 119, 176).

[Baa+22] Hendrik Baatz, Jonathan Granskog, Marios Papas, Fabrice Rousselle, and Jan Novák. “NeRF-Tex: Neural Reflectance Field Textures”. In: Computer Graphics Forum. Vol. 41. 6. Wiley Online Library. 2022, pp. 287–301 (cit. on pp. 104, 187).

[BSK23] Steve Bako, Pradeep Sen, and Anton Kaplanyan. “Deep Appearance Prefiltering”. In: ACM Transactions on Graphics (TOG) 42.2 (2023), pp. 1–23 (cit. on p. 105).

[BYK21] Wentao Bao, Qi Yu, and Yu Kong. “Evidential Deep Learning for Open Set Action Recognition”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 13349–13358 (cit. on p. 130).

[Bar+09] Connelly Barnes, Eli Shechtman, Adam Finkelstein, and Dan B Goldman. “PatchMatch: A Randomized Correspondence Algorithm for Structural Image Editing”. In: ACM Transactions on Graphics (TOG) 28.3 (2009), p. 24 (cit. on p. 74).

[Bar+15] Connelly Barnes, Fang-Lue Zhang, Liming Lou, Xian Wu, and Shi-Min Hu. “Patchtable: Efficient Patch Queries for Large Datasets and Applications”. In: ACM Transactions on Graphics (TOG) 34.4 (2015), pp. 1–10 (cit. on pp. 32, 74).

[Bar+21] Jonathan T. Barron, Ben Mildenhall, Matthew Tancik, Peter Hedman, Ricardo Martin-Brualla, and Pratul P. Srinivasan. “Mip-NeRF: A Multiscale Representation for Anti-Aliasing Neural Radiance Fields”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021 (cit. on p. 186).

[Ben+21] Saguy Benaim, Ron Mokady, Amit Bermano, and Lior Wolf. “Structural Analogy from a Single Image Pair”. In: Computer Graphics Forum. Vol. 40. 1. Wiley Online Library. 2021, pp. 249–265 (cit. on pp. 31, 33, 46, 47, 56, 57, 60, 75).

[Bén+13] Pierre Bénard, Forrester Cole, Michael Kass, Igor Mordatch, James Hegarty, Martin Sebastian Senn, Kurt Fleischer, Davide Pesare, and Katherine Breeden. “Stylizing Animation by Example”. In: ACM Transactions on Graphics (TOG) 32.4 (2013), pp. 1–12 (cit. on p. 32).

[Ben+20] Gregory Benton, Marc Finzi, Pavel Izmailov, and Andrew Gordon Wilson. “Learning Invariances in Neural Networks”. In: arXiv preprint arXiv:2010.11882 (2020) (cit. on p. 43).

[BJV17] Urs Bergmann, Nikolay Jetchev, and Roland Vollgraf. “Learning Texture Manifolds with the Periodic Spatial GAN”. In: International Conference on Machine Learning (ICML). 2017, pp. 469–477 (cit. on pp. 72, 74–76, 89–91, 95, 99).

[Ber+23] Hugo Bertiche, Niloy J Mitra, Kuldeep Kulkarni, Chun-Hao Paul Huang, Tuanfeng Y Wang, Meysam Madadi, Sergio Escalera, and Duygu Ceylan. “Blowing in the Wind: CycleNet for Human Cinemagraphs from Still Images”. In: arXiv preprint arXiv:2303.08639 (2023) (cit. on pp. 17, 259).

[Bha+03] Kiran S. Bhat, Christopher D. Twigg, Jessica K. Hodgins, Pradeep K. Khosla, Zoran Popovic, and Steven M. Seitz. “Estimating Cloth Simulation Parameters from Video”. In: Symposium on Computer Animation. The Eurographics Association, 2003 (cit. on pp. 156, 158).

[Bi+18] Wenyan Bi, Peiran Jin, Hendrikje Nienborg, and Bei Xiao. “Estimating Mechanical Properties of Cloth from Videos Using Dense Motion Trajectories: Human Psychophysics and Machine Learning”. In: Journal of vision 18.5 (2018), pp. 12–12 (cit. on p. 159).

[Bie20] Lukas Biewald. Experiment Tracking with Weights and Biases. Software available from wandb.com. 2020 (cit. on pp. 108, 176, 194).

[Bit+20] Benedikt Bitterli, Chris Wyman, Matt Pharr, Peter Shirley, Aaron Lefohn, and Wojciech Jarosz. “Spatiotemporal Reservoir Resampling for Real-time Ray Tracing with Dynamic Direct Lighting”. In: ACM Transactions on Graphics (Proceedings of SIGGRAPH) 39.4 (2020) (cit. on p. 184).

[Blu+16] Adrian Blumer, Jan Novák, Ralf Habel, Derek Nowrouzezahrai, and Wojciech Jarosz. “Reduced Aggregate Scattering Operators for Path Tracing”. In: Computer Graphics Forum (Proceedings of Pacific Graphics) 35.7 (2016), pp. 461–473 (cit. on p. 203).

[Bom+21] Rishi Bommasani, Drew A Hudson, Ehsan Adeli, Russ Altman, Simran Arora, Sydney von Arx, Michael S Bernstein, Jeannette Bohg, Antoine Bosselut, Emma Brunskill, et al. “On the Opportunities and Risks of Foundation Models”. In: (2021) (cit. on p. 159).

[Bos+20] Mark Boss, Varun Jampani, Kihwan Kim, Hendrik Lensch, and Jan Kautz. “Two-shot Spatially-Varying BRDF and Shape Estimation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 3982–3991 (cit. on p. 125).

[Bou+13] Katherine L Bouman, Bei Xiao, Peter Battaglia, and William T Freeman. “Estimating the Material Properties of Fabric from Video”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2013, pp. 1984–1991 (cit. on pp. 156, 159).

[BS12] Brent Burley and Walt Disney Animation Studios. “Physically-based Shading at Disney”. In: ACM SIGGRAPH. Vol. 2012. vol. 2012. 2012, pp. 1–7 (cit. on pp. 76, 127).

[Byl+22] Zoya Bylinskii, Laura Herman, Aaron Hertzmann, Stefanie Hutka, and Yile Zhang. “Towards Better User Studies in Computer Graphics and Vision”. In: arXiv preprint arXiv:2206.11461 (2022) (cit. on pp. 178, 182).

[CLA19] Carlos Castillo, Jorge López-Moreno, and Carlos Aliaga. “Recent Advances in Fabric Appearance Reproduction”. In: Computers & Graphics 84 (2019), pp. 103–121 (cit. on pp. 14, 17, 40, 256, 259).

[CHB21] Thomas Chambon, Eric Heitz, and Laurent Belcour. “Passing Multi-Channel Material Textures to a 3-Channel Loss”. In: ACM SIGGRAPH 2021 Talks. 2021, pp. 1–2 (cit. on pp. 64, 66, 67, 128).

[Cha+19] Caroline Chan, Shiry Ginosar, Tinghui Zhou, and Alexei A. Efros. “Everybody Dance Now”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2019 (cit. on p. 85).

[Che+20a] Chengqian Che, Fujun Luan, Shuang Zhao, Kavita Bala, and Ioannis Gkioulekas. “Towards Learning-based Inverse Subsurface Scattering”. In: 2020 IEEE International Conference on Computational Photography (ICCP). IEEE. 2020, pp. 1–12 (cit. on p. 203).

[Che+22] Anpei Chen, Zexiang Xu, Andreas Geiger, Jingyi Yu, and Hao Su. “Tensorf: Tensorial Radiance Fields”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2022, pp. 333–350 (cit. on pp. 203, 212).

[Che+17a] Dongdong Chen, Jing Liao, Lu Yuan, Nenghai Yu, and Gang Hua. “Coherent Online Video Style Transfer”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2017, pp. 1105–1114 (cit. on p. 32).

[Che+17b] Dongdong Chen, Lu Yuan, Jing Liao, Nenghai Yu, and Gang Hua. “Stylebank: An Explicit Representation for Neural Image Style Transfer”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 1897–1906 (cit. on p. 32).

[Che+20b] Ting Chen, Simon Kornblith, Mohammad Norouzi, and Geoffrey Hinton. “A Simple Framework for Contrastive Learning of Visual Representations”. In: International Conference on Machine Learning (ICML). PMLR. 2020, pp. 1597–1607 (cit. on pp. 37, 45, 159).

[CXH21] Xinlei Chen, Saining Xie, and Kaiming He. “An Empirical Study of Training Self-supervised Vision Transformers”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2021, pp. 9640–9649 (cit. on p. 159).

[Che+18] Yuanqin Chen, Qian Zhang, Yaping Wu, Bo Liu, Meiyun Wang, and Yusong Lin. “Fine-tuning ResNet for Breast Cancer Classification from mammography”. In: The International Conference on Healthcare Science and Engineering. Springer. 2018, pp. 83–96 (cit. on p. 159).

[CNN22] Zhe Chen, Shohei Nobuhara, and Ko Nishino. “Invertible Neural BRDF for Object Inverse Rendering”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 44.12 (2022), pp. 9380–9395 (cit. on pp. 186, 203).

[Cho+18] Yunjey Choi, Minje Choi, Munyoung Kim, Jung-Woo Ha, Sunghun Kim, and Jaegul Choo. “Stargan: Unified Generative Adversarial Networks for Multi-domain Image-to-image Translation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 8789–8797 (cit. on p. 77).

[Cla+90] Timothy G Clapp, Hong Peng, Tushar K Ghosh, and Jeffrey W Eischen. “Indirect Measurement of the Moment-curvature Relationship for Fabrics”. In: Textile Research Journal 60.9 (1990), pp. 525–533 (cit. on pp. 156, 158).

[Cla+21] Petrik Clarberg, Wojciech Jarosz, Tomas Akenine-Möller, and Henrik Wann Jensen. “Wavelet Importance Sampling: Efficiently Evaluating Products of Complex Functions”. In: ACM Transactions on Graphics (Proceedings of SIGGRAPH) 24.3 (2021), pp. 3520–3528 (cit. on pp. 185, 189, 195, 199, 200).

[CTT17] David Clyde, Joseph Teran, and Rasmus Tamstorf. “Modeling and Data-driven Parameter Estimation for Woven Fabrics”. In: Proceedings of the ACM SIGGRAPH/Eurographics Symposium on Computer Animation. 2017, pp. 1–11 (cit. on p. 158).

[Coh+03] Michael F Cohen, Jonathan Shade, Stefan Hiller, and Oliver Deussen. “Wang Tiles for Image and Texture Generation”. In: ACM Transactions on Graphics (TOG) 22.3 (2003), pp. 287–294 (cit. on p. 72).

[CT82] Robert L Cook and Kenneth E. Torrance. “A Reflectance Model for Computer Graphics”. In: ACM Transactions on Graphics (TOG) 1.1 (1982), pp. 7–24 (cit. on p. 48).

[Dan01] Kristin J Dana. “BRDF/BTF Measurement Device”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Vol. 2. IEEE. 2001, pp. 460–466 (cit. on p. 103).

[Dan+99] Kristin J Dana, Bram Van Ginneken, Shree K Nayar, and Jan J Koenderink. “Reflectance and Texture of Real-World Surfaces”. In: ACM Transactions On Graphics (TOG) 18.1 (1999), pp. 1–34 (cit. on p. 33).

[Dar+12] Soheil Darabi, Eli Shechtman, Connelly Barnes, Dan B Goldman, and Pradeep Sen. “Image Melding: Combining Inconsistent Images Using Patch-Based Synthesis”. In: ACM Transactions on Graphics (TOG) 31.4 (2012), pp. 1–10 (cit. on p. 74).

[Dav+15] Abe Davis, Katherine L Bouman, Justin G Chen, Michael Rubinstein, Fredo Durand, and William T Freeman. “Visual Vibrometry: Estimating Material Properties from Small Motion in Video”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2015, pp. 5335–5343 (cit. on p. 159).

[De 97] Jeremy S De Bonet. “Multiresolution Sampling Procedure for Analysis and Synthesis of Texture Images”. In: Proceedings of the 24th Annual cCnference on Computer Graphics and Interactive Techniques. 1997, pp. 361–368 (cit. on p. 74).

[Dek+15] Tali Dekel, Tomer Michaeli, Michal Irani, and William T. Freeman. “Revealing and Modifying Non-Local Variations in a Single Image”. In: ACM Transactions on Graphics (TOG) (2015) (cit. on p. 93).

[DH19] Thomas Deliot and Eric Heitz. “Procedural Stochastic Textures by Tiling and Blending”. In: GPU Zen 2 (2019) (cit. on pp. 72, 74, 89–91, 95, 99).

[Den+09] Jia Deng, Wei Dong, Richard Socher, Li-Jia Li, Kai Li, and Li Fei-Fei. “Imagenet: A Large-scale Hierarchical Image Database”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2009, pp. 248–255 (cit. on pp. 32, 164, 166).

[Des+18] Valentin Deschaintre, Miika Aittala, Fredo Durand, George Drettakis, and Adrien Bousseau. “Single-image SVBRDF Capture with a Rendering-Aware Deep Network”. In: ACM Transactions on Graphics (TOG) 37.4 (2018) (cit. on pp. 17, 34, 39, 54, 77, 102, 123, 125, 127, 258).

[Des+19] Valentin Deschaintre, Miika Aittala, Frédo Durand, George Drettakis, and Adrien Bousseau. “Flexible SVBRDF Capture with a Multi-Image Deep Network”. In: Computer Graphics Forum. Vol. 38. 4. Wiley Online Library. 2019, pp. 1–13 (cit. on pp. 34, 42, 72, 76, 102, 125, 141).

[DDB20] Valentin Deschaintre, George Drettakis, and Adrien Bousseau. “Guided Fine-Tuning for Large-Scale Material Transfer”. In: Computer Graphics Forum. Vol. 39. 4. Wiley Online Library. 2020, pp. 91–105 (cit. on pp. 35, 38–40, 47, 49, 50, 63, 72, 76, 105, 125, 142).

[DLG21] Valentin Deschaintre, Yiming Lin, and Abhijeet Ghosh. “Deep Polarization Imaging for 3D Shape and SVBRDF Acquisition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 15567–15576 (cit. on pp. 64, 67, 69, 127, 141).

[Dia+20] Foivos I Diakogiannis, François Waldner, Peter Caccetta, and Chen Wu. “ResUNet-a: A Deep Learning Framework for Semantic Segmentation of Remotely Sensed Data”. In: ISPRS Journal of Photogrammetry and Remote Sensing 162 (2020), pp. 94–114 (cit. on pp. 64, 66, 69, 108, 127, 141).

[Dia+15] Olga Diamanti, Connelly Barnes, Sylvain Paris, Eli Shechtman, and Olga Sorkine-Hornung. “Synthesis of Complex Image Appearance from Limited Exemplars”. In: ACM Transactions on Graphics (TOG) 34.2 (2015), pp. 1–14 (cit. on pp. 34, 104).

[Din+20] Keyan Ding, Kede Ma, Shiqi Wang, and Eero P Simoncelli. “Image Quality Assessment: Unifying Structure and Texture Similarity”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence (2020) (cit. on p. 160).

[DKB14] Laurent Dinh, David Krueger, and Yoshua Bengio. “Nice: Non-linear Independent Components Estimation”. In: arXiv preprint arXiv:1410.8516 (2014) (cit. on p. 191).

[DSB16] Laurent Dinh, Jascha Sohl-Dickstein, and Samy Bengio. “Density Estimation Using Real NVP”. In: arXiv preprint arXiv:1605.08803 (2016) (cit. on p. 191).

[Dod+22] Ana Dodik, Silvia Sellán, Theodore Kim, and Amanda Phillips. “Sex and Gender in the Computer Graphics Research Literature”. In: arXiv preprint arXiv:2206.00480 (2022) (cit. on pp. 178, 179, 182).

[DZP07] Weiming Dong, Ning Zhou, and Jean-Claude Paul. “Optimized Tile-Based Texture Synthesis”. In: Proceedings of Graphics Interface 2007. 2007, pp. 249–256 (cit. on p. 74).

[Don19] Yue Dong. “Deep Appearance Modeling: A Survey”. In: Visual Informatics 3.2 (2019), pp. 59–68 (cit. on pp. 34, 53).

[Don+10] Yue Dong, Jiaping Wang, Xin Tong, John Snyder, Yanxiang Lan, Moshe Ben-Ezra, and Baining Guo. “Manifold Bootstrapping for SVBRDF Capture”. In: ACM Transactions on Graphics (TOG) 29.4 (2010), pp. 1–10 (cit. on p. 34).

[DB16] Alexey Dosovitskiy and Thomas Brox. “Generating images with Perceptual Similarity Metrics based on Deep Networks”. In: Advances in Neural Information Processing Systems. 2016, pp. 658–666 (cit. on p. 74).

[DTD21] Emilien Dupont, Yee Whye Teh, and Arnaud Doucet. “Generative Models as Distributions of Functions”. In: arXiv preprint arXiv:2102.04776 (2021) (cit. on p. 75).

[Dur+19a] Conor Durkan, Artur Bekasov, Iain Murray, and George Papamakarios. “Neural Spline Flows”. In: Advances in Neural Information Processing Systems. Vol. 32. 2019 (cit. on pp. 185, 191, 192, 194, 196, 197).

[Dur+19b] Conor Durkan, Artur Bekasovs, Iain Murray, and Georgios Papamakarios. “Cubic-Spline Flows”. In: Workshop on Invertible Neural Nets and Normalizing Flows: ICML 2019. 2019 (cit. on p. 191).

[Dye+18] Joanne Dyer, Diego Tamburini, Elisabeth R O’Connell, and Anna Harrison. “A Multispectral Imaging Approach Integrated into the Study of Late Antique textiles from Egypt”. In: PLoS One 13.10 (2018), e0204699 (cit. on p. 213).

[EF01] Alexei A Efros and William T Freeman. “Image Quilting for Texture Synthesis and Transfer”. In: Proceedings of the 28th annual conference on Computer Graphics and Interactive Techniques. 2001, pp. 341–346 (cit. on pp. 72, 74).

[EL99] Alexei A Efros and Thomas K Leung. “Texture Synthesis by Non-parametric Sampling”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Vol. 2. IEEE. 1999, pp. 1033–1038 (cit. on p. 74).

[EM17] Michael Elad and Peyman Milanfar. “Style Transfer via Texture Synthesis”. In: IEEE Transactions on Image Processing 26.5 (2017), pp. 2338–2351 (cit. on pp. 34, 104).

[EUD18] Stefan Elfwing, Eiji Uchibe, and Kenji Doya. “Sigmoid-Weighted Linear Units for Neural Network Function Approximation in Reinforcement Learning”. In: Neural Networks 107 (2018), pp. 3–11 (cit. on pp. 141, 144).

[End+16] Yuki Endo, Satoshi Iizuka, Yoshihiro Kanamori, and Jun Mitani. “Deepprop: Extracting Deep Features from a Single Image for Edit Propagation”. In: Computer Graphics Forum. Vol. 35. 2. Wiley Online Library. 2016, pp. 189–201 (cit. on p. 33).

[Fan+22] Jiahui Fan, Beibei Wang, Miloš Hašan, Jian Yang, and Ling-Qi Yan. “Neural Layered BRDFs”. In: Proceedings of SIGGRAPH 2022. 2022 (cit. on pp. 116, 187).

[Fen+22] Xudong Feng, Wenchao Huang, Weiwei Xu, and Huamin Wang. “Learning-Based Bending Stiffness Parameter Estimation by a Drape Tester”. In: ACM Transactions on Graphics (TOG) 41.6 (2022), pp. 1–16 (cit. on p. 159).

[Fil+18] J. Filip, M. Kolafová, M. Havlıček, R. Vávra, M. Haindl, and Rushmeier H. “Evaluating Physical and Rendered Material Appearance”. In: The Visual Computer (Computer Graphics International 2018) (2018) (cit. on p. 109).

[FH08] Jiřı Filip and Michal Haindl. “Bidirectional Texture Function Modeling: A State of the Art Survey”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 31.11 (2008), pp. 1921–1940 (cit. on p. 103).

[FR22] Michael Fischer and Tobias Ritschel. “Plateau-Reduced Differentiable Path Tracing”. In: arXiv preprint arXiv:2211.17263 (2022) (cit. on p. 202).

[Fiš+16] Jakub Fišer, Ondřej Jamriška, Michal Lukáč, Eli Shechtman, Paul Asente, Jingwan Lu, and Daniel Sy`kora. “StyLit: Illumination-Guided Example-Based Stylization of 3D Renderings”. In: ACM Transactions on Graphics (TOG) 35.4 (2016), pp. 1–11 (cit. on p. 33).

[Fri+22] Sara Fridovich-Keil, Alex Yu, Matthew Tancik, Qinhong Chen, Benjamin Recht, and Angjoo Kanazawa. “Plenoxels: Radiance Fields Without Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 5501–5510 (cit. on pp. 187, 212).

[FAW19] Anna Frühstück, Ibraheem Alhashim, and Peter Wonka. “TileGAN: Synthesis of Large-Scale Non-Homogeneous Textures”. In: ACM Transactions on Graphics (TOG) 38.4 (Apr. 2019) (cit. on pp. 34, 72, 74, 79–81, 104).

[Fu+20] Ruigang Fu, Qingyong Hu, Xiaohu Dong, Yulan Guo, Yinghui Gao, and Biao Li. “Axiom-based Grad-CAM: Towards Accurate Visualization and Explanation of CNNs”. In: 31st British Machine Vision Conference 2020, BMVC 2020, BMVA Press, 2020 (cit. on pp. 167, 168).

[GG16] Yarin Gal and Zoubin Ghahramani. “Dropout as a Bayesian Approximation: Representing Model Uncertainty in Deep Learning”. In: International Conference on Machine Learning (ICML). PMLR. 2016, pp. 1050–1059 (cit. on pp. 123, 126, 130, 141).

[Gal+12] Bruno Galerne, Ares Lagae, Sylvain Lefebvre, and George Drettakis. “Gabor Noise by Example”. In: ACM Transactions on Graphics (TOG) 31.4 (2012), pp. 1–9 (cit. on p. 72).

[Gao+21] Chen Gao, Ayush Saraf, Johannes Kopf, and Jia-Bin Huang. “Dynamic View Synthesis From Dynamic Monocular Video”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 5712–5721 (cit. on p. 186).

[Gao+19] Duan Gao, Xiao Li, Yue Dong, Pieter Peers, Kun Xu, and Xin Tong. “Deep Inverse Rendering for High-Resolution SVBRDF Estimation from an Arbitrary Number of Images”. In: ACM Transactions on Graphics (TOG) 38.4 (2019), pp. 1–15 (cit. on pp. 49, 59, 61, 125, 137, 138, 152, 153).

[GMX22] Duan Gao, Haoyuan Mu, and Kun Xu. “Neural Global Illumination: Interactive In-direct Illumination Prediction under Dynamic Area Lights”. In: IEEE Transactions on Visualization and Computer Graphics (2022) (cit. on pp. 105, 186, 203).

[Gar+14] Elena Garces, Aseem Agarwala, Diego Gutierrez, and Aaron Hertzmann. “A Similarity Measure for Illustration Style”. In: ACM Transactions on Graphics (TOG) 33.4 (2014), pp. 1–9 (cit. on pp. 157, 160, 170, 178, 180).

[Gar+23] Elena Garces, Victor Arellano, Carlos Rodriguez-Pardo, David Pascual-Hernandez, Sergio Suja, and Jorge Lopez-Moreno. “Towards Material Digitization with a Dual-scale Optical System”. In: ACM Transactions on Graphics (TOG) (2023) (cit. on pp. 15, 16, 24, 102, 109, 269).

[Gar+22] Elena Garces, Carlos Rodriguez-Pardo, Dan Casas, and Jorge Lopez-Moreno. “A Survey on Intrinsic Images: Delving Deep into Lambert and Beyond”. In: International Journal of Computer Vision (2022) (cit. on pp. 3, 4, 6, 7, 12, 22, 23, 127, 128, 141, 268).

[GES22] James Gardner, Bernhard Egger, and William Alfred Peter Smith. “Rotation-Equivariant Conditional Spherical Neural Fields for Learning a Natural Illumination Prior”. In: Advances in Neural Information Processing Systems. 2022 (cit. on pp. 19, 185, 186, 193, 196, 199, 202, 203, 260).

[Gar+19] Marc-André Gardner, Yannick Hold-Geoffroy, Kalyan Sunkavalli, Christian Gagné, and Jean-François Lalonde. “Deep Parametric Indoor Lighting Estimation”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2019, pp. 7175–7183 (cit. on p. 203).

[GEB15a] Leon Gatys, Alexander S Ecker, and Matthias Bethge. “Texture Synthesis using Convolutional Neural Networks”. In: Advances in Neural Information Processing Systems. 2015, pp. 262–270 (cit. on pp. 74, 95, 128).

[GEB15b] Leon A Gatys, Alexander S Ecker, and Matthias Bethge. “A Neural Algorithm of Artistic Style”. In: arXiv preprint arXiv:1508.06576 (2015) (cit. on pp. 32, 47, 48, 124, 128).

[GEB16a] Leon A Gatys, Alexander S Ecker, and Matthias Bethge. “Image Style Transfer using Convolutional Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2016, pp. 2414–2423 (cit. on pp. 74, 160).

[Gat+17] Leon A Gatys, Alexander S Ecker, Matthias Bethge, Aaron Hertzmann, and Eli Shechtman. “Controlling Perceptual Factors in Neural Style Transfer”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 3985–3993 (cit. on pp. 32, 74).

[GEB16b] Leon A. Gatys, Alexander S. Ecker, and Matthias Bethge. “Image Style Transfer Using Convolutional Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vol. 2016-Decem. June 2016, pp. 2414–2423 (cit. on p. 79).

[Gau+22] Alban Gauthier, Robin Faury, Jérémy Levallois, Theo Thonat, Jean-Marc Thiery, and Tamy Boubekeur. “MIPNet: Neural Normal-to-Anisotropic-Roughness MIP mapping”. In: ACM Transactions on Graphics (TOG) 41.6 (2022), pp. 1–12 (cit. on p. 105).

[Gaw+21] Jakob Gawlikowski, Cedrique Rovile Njieutcheu Tassi, Mohsin Ali, Jongseok Lee, Matthias Humt, Jianxiang Feng, Anna Kruspe, Rudolph Triebel, Peter Jung, Ribana Roscher, et al. “A Survey of Uncertainty in Deep Neural Networks”. In: arXiv preprint arXiv:2107.03342 (2021) (cit. on p. 126).

[Geo+18] Iliyan Georgiev, Thiago Ize, Mike Farnsworth, Ramón Montoya-Vozmediano, Alan King, Brecht Van Lommel, Angel Jimenez, Oscar Anson, Shinji Ogaki, Eric Johnston, et al. “Arnold: A Brute-force Production Path Tracer”. In: ACM Transactions on Graphics (TOG) 37.3 (2018), pp. 1–12 (cit. on p. 48).

[Gil+14] Guillaume Gilet, Basile Sauvage, Kenneth Vanhoey, Jean-Michel Dischler, and Djamchid Ghazanfarpour. “Local Random-phase Noise for Procedural Texturing”. In: ACM Transactions on Graphics (TOG) 33.6 (2014), pp. 1–11 (cit. on p. 72).

[Goo+14a] Ian Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. “Generative Adversarial Nets”. In: Advances in Neural Information Processing Systems. 2014, pp. 2672–2680 (cit. on p. 39).

[Goo+14b] Ian J. Goodfellow, Jean Pouget-Abadie, Mehdi Mirza, Bing Xu, David Warde-Farley, Sherjil Ozair, Aaron Courville, and Yoshua Bengio. “Generative Adversarial Nets”. In: Advances in Neural Information Processing Systems. Vol. 3. January. Neural information processing systems foundation, June 2014, pp. 2672–2680 (cit. on p. 79).

[Gri+03] Eitan Grinspun, Anil N Hirani, Mathieu Desbrun, and Peter Schröder. “Discrete Shells”. In: Proceedings of the 2003 ACM SIGGRAPH/Eurographics symposium on Computer animation. Citeseer. 2003, pp. 62–67 (cit. on pp. 161, 162).

[Gro+20] Aditya Grover, Christopher Chute, Rui Shu, Zhangjie Cao, and Stefano Ermon. “Align-flow: Cycle Consistent Learning from Multiple Domains via Normalizing Flows”. In: Proceedings of the AAAI Conference on Artificial Intelligence. Vol. 34. 04. 2020, pp. 4028–4035 (cit. on p. 186).

[Gu+18] Shuyang Gu, Congliang Chen, Jing Liao, and Lu Yuan. “Arbitrary Style Transfer with Deep Feature Reshuffle”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 8222–8231 (cit. on p. 32).

[Gua+16] Darya Guarnera, Giuseppe Claudio Guarnera, Abhijeet Ghosh, Cornelia Denk, and Mashhuda Glencross. “BRDF Representation and Acquisition”. In: Computer Graphics Forum. Vol. 35. 2. Wiley Online Library. 2016, pp. 625–650 (cit. on p. 34).

[Gua+19] Giuseppe Claudio Guarnera, Dar’ya Guarnera, Gregory J Ward, Mashhuda Glencross, and Ian Hall. “BxDF Material Acquisition, Representation, and Rendering for VR and Design”. In: SIGGRAPH Asia 2019 Courses. 2019, pp. 1–21 (cit. on p. 3).

[Gue+20] P Guehl, R Allègre, J-M Dischler, B Benes, and E Galin. “Semi-Procedural Textures Using Point Process Texture Basis Functions”. In: Computer Graphics Forum. Vol. 39. 4. Wiley Online Library. 2020, pp. 159–171 (cit. on pp. 34, 72, 92, 104).

[Gui+17] Geoffrey Guingo, Basile Sauvage, Jean-Michel Dischler, and Marie-Paule Cani. “Bilayer Textures: A Model for Synthesis and Deformation of Composite Textures”. In: Computer Graphics Forum. Vol. 36. 4. 2017, pp. 111–122 (cit. on p. 72).

[Guo+21] Jie Guo, Shuichang Lai, Chengzhi Tao, Yuelong Cai, Lei Wang, Yanwen Guo, and Ling-Qi Yan. “Highlight-Aware Two-Stream Network for Single-Image SVBRDF Acquisition”. In: ACM Transactions on Graphics (TOG) 40.4 (2021), pp. 1–14 (cit. on pp. 123, 125).

[Guo+20a] Yu Guo, Miloš Hašan, Lingqi Yan, and Shuang Zhao. “A Bayesian Inference Framework for Procedural Material Parameter Estimation”. In: Computer Graphics Forum. Vol. 39. 7. Wiley Online Library. 2020, pp. 255–266 (cit. on pp. 125, 126).

[Guo+20b] Yu Guo, Cameron Smith, Miloš Hašan, Kalyan Sunkavalli, and Shuang Zhao. “MaterialGAN: Reflectance Capture using a Generative SVBRDF Model”. In: ACM Transactions on Graphics (TOG) 39.6 (2020), 254:1–254:13 (cit. on pp. 17, 35, 42, 72, 76, 102, 123, 125, 258).

[Guo+19] Yunhui Guo, Honghui Shi, Abhishek Kumar, Kristen Grauman, Tajana Rosing, and Rogerio Feris. “Spottune: Transfer Learning through Adaptive Fine-tuning”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4805–4814 (cit. on p. 159).

[HDL16] David Ha, Andrew Dai, and Quoc V Le. “Hypernetworks”. In: arXiv preprint arXiv:1609.09106 (2016) (cit. on p. 75).

[HMV09] Simon Haegler, Pascal Müller, and Luc Van Gool. “Procedural Modeling for Digital Cultural Heritage”. In: EURASIP Journal on Image and Video Processing 2009 (2009), pp. 1–1 (cit. on p. 213).

[Hai12] Vávra R. Haindl M. Filip J. “Digital Material Appearance: the Curse of Tera-Bytes”. In: ERCIM News 90 (2012), pp. 49–50 (cit. on pp. 108, 109, 111).

[Ham+21] Hendrik Hameeuw, Godelieve Watteeuw, Bruno Vandermeulen, and Marc Proesmans. “The Painted Panels of the Early Sixteenth Century Mechelen Enclosed Gardens Art Technical Examination with the Photometric Stereo: White Light and Multispectral Microdomes”. In: Papers Presented at the Twentieth Symposium for the Study of Underdrawing and Technology in Painting held in Mechelen and Leuven, 11-13 January 2017. Vol. 20. Peeters Publishers; Leuven. 2021, pp. 165–175 (cit. on p. 213).

[Han+22a] Jiyeon Han, Hwanil Choi, Yunjey Choi, Junho Kim, Jung-Woo Ha, and Jaesik Choi. “Rarity Score: A New Metric to Evaluate the Uncommonness of Synthesized Images”. In: arXiv preprint arXiv:2206.08549 (2022) (cit. on p. 128).

[Han+22b] Kai Han, Yunhe Wang, Hanting Chen, Xinghao Chen, Jianyuan Guo, Zhenhua Liu, Yehui Tang, An Xiao, Chunjing Xu, Yixing Xu, et al. “A Survey on Vision Transformer”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence (2022) (cit. on p. 164).

[Haq+23] Ayaan Haque, Matthew Tancik, Alexei Efros, Aleksander Holynski, and Angjoo Kanazawa. “Instruct-NeRF2NeRF: Editing 3D Scenes with Instructions”. In: arXiv preprint arXiv:2303.12789 (2023) (cit. on p. 116).

[HHM22] Jon Hasselgren, Nikolai Hofmann, and Jacob Munkberg. “Shape, Light & Material Decomposition from Images Using Monte Carlo Rendering and Denoising”. In: arXiv preprint arXiv:2206.03380 (2022) (cit. on p. 203).

[HFM10] Vlastimil Havran, Jirı Filip, and Karol Myszkowski. “Bidirectional Texture Function Compression Based on Multi-level Vector Quantization”. In: Computer Graphics Forum. Vol. 29. 1. Wiley Online Library. 2010, pp. 175–190 (cit. on p. 103).

[He+22] Kaiming He, Xinlei Chen, Saining Xie, Yanghao Li, Piotr Dollár, and Ross Girshick. “Masked Autoencoders are Scalable Vision Learners”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 16000–16009 (cit. on p. 159).

[He+16] Kaiming He, Xiangyu Zhang, Shaoqing Ren, and Jian Sun. “Deep Residual Learning for Image Recognition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vol. 2016-Decem. 2016, pp. 770–778 (cit. on pp. 64, 66, 69, 78, 94, 108, 109, 127, 141, 164, 176).

[He+18] Mingming He, Dongdong Chen, Jing Liao, Pedro V Sander, and Lu Yuan. “Deep Exemplar-Based Colorization”. In: ACM Transactions on Graphics (TOG) 37.4 (2018), pp. 1–16 (cit. on p. 33).

[He+19] Mingming He, Jing Liao, Dongdong Chen, Lu Yuan, and Pedro V Sander. “Progressive Color Transfer with Dense Semantic Correspondences”. In: ACM Transactions on Graphics (TOG) 38.2 (2019), pp. 1–18 (cit. on p. 33).

[HB95] David J Heeger and James R Bergen. “Pyramid-based Texture Analysis/Synthesis”. In: Proceedings of the 22nd Annual Conference on Computer Graphics and Interactive Techniques. 1995, pp. 229–238 (cit. on p. 74).

[Hei20] Eric Heitz. “Can’t Invert the CDF? The Triangle-Cut Parameterization of the Region under the Curve”. In: Computer Graphics Forum 39.4 (2020), pp. 121–132 (cit. on p. 185).

[HN18] Eric Heitz and Fabrice Neyret. “High-Performance By-Example Noise Using a Histogram-Preserving Blending Operator”. In: Proceedings of the ACM on Computer Graphics and Interactive Techniques 1.2 (2018), pp. 1–25 (cit. on p. 72).

[Hei+21] Eric Heitz, Kenneth Vanhoey, Thomas Chambon, and Laurent Belcour. “A Sliced Wasserstein Loss for Neural Texture Synthesis”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 9412–9420 (cit. on pp. 75, 93).

[Hel+20] Leonhard Helminger, Abdelaziz Djelouah, Markus Gross, and Christopher Schroers. “Lossy Image Compression with Normalizing Flows”. In: arXiv preprint arXiv:2008.10486 (2020) (cit. on p. 186).

[HG16] Dan Hendrycks and Kevin Gimpel. “Gaussian Error Linear Units (GELUs)”. In: arXiv preprint arXiv:1606.08415 (2016) (cit. on pp. 109, 119).

[Hen+21] Philipp Henzler, Valentin Deschaintre, Niloy J Mitra, and Tobias Ritschel. “Generative Modelling of BRDF Textures from Flash Images”. In: ACM Transactions on Graphics (Proc. SIGGRAPH Asia) 40.6 (2021) (cit. on pp. 17, 76, 102, 123, 124, 137, 138, 143, 151–153, 258).

[Her+20] Amir Hertz, Rana Hanocka, Raja Giryes, and Daniel Cohen-Or. “Deep Geometric Texture Synthesis”. In: ACM Transactions on Graphics (TOG) 39.4 (2020), pp. 108–1 (cit. on p. 72).

[Her+01] Aaron Hertzmann, Charles E Jacobs, Nuria Oliver, Brian Curless, and David H Salesin. “Image Analogies”. In: Proceedings of the 28th annual conference on Computer graphics and interactive techniques. 2001, pp. 327–340 (cit. on pp. 31, 32, 46, 56, 57, 60).

[HS05] Aaron Hertzmann and Steven M Seitz. “Example-based Photometric Stereo: Shape Reconstruction with General, Varying BRDFs”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 27.8 (2005), pp. 1254–1264 (cit. on p. 34).

[HS03] Aaron Hertzmann and Steven M Seitz. “Shape and Materials by example: A Photometric Stereo Approach”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). Vol. 1. IEEE. 2003, pp. I–I (cit. on p. 34).

[Heu+17] Martin Heusel, Hubert Ramsauer, Thomas Unterthiner, Bernhard Nessler, and Sepp Hochreiter. “Gans Trained by a Two Time-Scale Update Rule Converge to a Local Nash Equilibrium”. In: arXiv preprint arXiv:1706.08500 (2017) (cit. on p. 85).

[Hil+15] Stephen Hill, Stephen McAuley, Brent Burley, et al. “Physically Based Shading in Theory and Practice”. In: ACM SIGGRAPH 2015 Courses. SIGGRAPH ’15. Los Angeles, California: Association for Computing Machinery, 2015 (cit. on p. 11).

[HS06] Geoffrey E Hinton and Ruslan R Salakhutdinov. “Reducing the Dimensionality of Data with Neural Networks”. In: science 313.5786 (2006), pp. 504–507 (cit. on p. 103).

[Hin+21] Tobias Hinz, Matthew Fisher, Oliver Wang, and Stefan Wermter. “Improved Techniques for Training Single-Image GANs”. In: Proceedings - 2021 IEEE Winter Conference on Applications of Computer Vision, WACV 2021. Jan. 2021, pp. 1300–1309 (cit. on p. 75).

[HJA20] Jonathan Ho, Ajay Jain, and Pieter Abbeel. “Denoising Diffusion Probabilistic Models”. In: Advances in Neural Information Processing Systems 33 (2020), pp. 6840–6851 (cit. on pp. 127, 209).

[HS13] Shahera Hossain and Seiichi Serikawa. “Texture Databases–A Comprehensive Survey”. In: Pattern Recognition Letters 34.15 (2013), pp. 2007–2022 (cit. on p. 109).

[Hu+20] Bingyang Hu, Jie Guo, Yanjun Chen, Mengtian Li, and Yanwen Guo. “DeepBRDF: A Deep Representation for Manipulating Measured BRDF”. In: Computer Graphics Forum. Vol. 39. 2. Wiley Online Library. 2020, pp. 157–166 (cit. on pp. 105, 203).

[Hu+22a] Dongting Hu, Liuhua Peng, Tingjin Chu, Xiaoxing Zhang, Yinian Mao, Howard Bondell, and Mingming Gong. “Uncertainty Quantification in Depth Estimation via Constrained Ordinal Regression”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2022 (cit. on pp. 126, 130).

[HSS18] Jie Hu, Li Shen, and Gang Sun. “Squeeze-and-excitation Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 7132–7141 (cit. on p. 164).

[HXP20] Wei Hu, Lechao Xiao, and Jeffrey Pennington. “Provable Benefit of Orthogonal Initialization in Optimizing Deep Linear Networks”. In: arXiv preprint arXiv:2001.05992 (2020) (cit. on p. 141).

[HDR19] Yiwei Hu, Julie Dorsey, and Holly Rushmeier. “A Novel Framework for Inverse Procedural Texture Modeling”. In: ACM Transactions on Graphics (TOG) 38.6 (2019), pp. 1–14 (cit. on p. 72).

[Hu+22b] Yiwei Hu, Miloš Hašan, Paul Guerrero, Holly Rushmeier, and Valentin Deschaintre. “Controlling Material Appearance by Examples”. In: Computer Graphics Forum. Vol. 41. 4. Wiley Online Library. 2022, pp. 117–128 (cit. on p. 105).

[Hu+22c] Yiwei Hu, Chengan He, Valentin Deschaintre, Julie Dorsey, and Holly Rushmeier. “An Inverse Procedural Modeling Pipeline for SVBRDF Maps”. In: ACM Transactions on Graphics (TOG) 41.2 (2022), pp. 1–17 (cit. on p. 125).

[Hu+19] Yuanming Hu, Luke Anderson, Tzu-Mao Li, Qi Sun, Nathan Carr, Jonathan Ragan-Kelley, and Fredo Durand. “Diff Taichi: Differentiable Programming for Physical Simulation”. In: International Conference on Learning Representations (ICLR). 2019 (cit. on p. 158).

[Hua+18a] Chin-Wei Huang, David Krueger, Alexandre Lacoste, and Aaron Courville. “Neural Autoregressive Flows”. In: International Conference on Machine Learning (ICML). PMLR. 2018, pp. 2078–2087 (cit. on p. 191).

[HB17] Xun Huang and Serge Belongie. “Arbitrary Style Transfer in Real-time with Adaptive Instance Normalization”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2017, pp. 1501–1510 (cit. on p. 32).

[Hua+18b] Xun Huang, Ming-Yu Liu, Serge Belongie, and Jan Kautz. “Multimodal Unsupervised Image-to-image Translation”. In: Proceedings of the European Conference on Computer Vision (ECCV). Sept. 2018 (cit. on p. 85).

[Hua+19] Zhiyuan Huang, Mansur Arief, Henry Lam, and Ding Zhao. “Evaluation Uncertainty in Data-Driven Self-Driving Testing”. In: 2019 IEEE Intelligent Transportation Systems Conference (ITSC). IEEE. 2019, pp. 1902–1907 (cit. on p. 126).

[HEW17] Markus Huber, Bernhard Eberhardt, and Daniel Weiskopf. “Cloth Animation Retrieval Using a Motion-shape Signature”. In: IEEE Computer Graphics and Applications 37.6 (2017), pp. 52–64 (cit. on p. 159).

[HZ21] Drew A Hudson and C Lawrence Zitnick. “Generative Adversarial Transformers”. In: arXiv preprint arXiv:2103.01209 (2021) (cit. on p. 93).

[Hwa+22] Inseung Hwang, Daniel S Jeon, Adolfo Muñoz, Diego Gutierrez, Xin Tong, and Min H Kim. “Sparse Ellipsometry: Portable Acquisition of Polarimetric SVBRDF and Shape with Unstructured Flash Photography”. In: ACM Transactions on Graphics (TOG) 41.4 (2022), pp. 1–14 (cit. on p. 129).

[Ike81] Katsushi Ikeuchi. “Determining Surface Orientations of Specular Surfaces by Using the Photometric Stereo Method”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 6 (1981), pp. 661–669 (cit. on pp. 34, 40, 50, 84, 112, 151).

[IS15] Sergey Ioffe and Christian Szegedy. “Batch Normalization: Accelerating Deep Network Training by Reducing Internal Covariate Shift”. In: International Conference on Machine Learning (ICML). PMLR. 2015, pp. 448–456 (cit. on pp. 54, 64, 194).

[Iso+17] Phillip Isola, Jun-Yan Zhu, Tinghui Zhou, and Alexei A. Efros. “Image-to-Image Translation with Conditional Adversarial Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2017, pp. 5967–5976 (cit. on pp. 38, 39, 54, 56, 78, 79, 94, 107, 128, 132, 193).

[Jak+22] Wenzel Jakob, Sébastien Speierer, Nicolas Roussel, Merlin Nimier-David, Delio Vicini, Tizian Zeltner, Baptiste Nicolet, Miguel Crespo, Vincent Leroy, and Ziyi Zhang. Mitsuba 3 Renderer. Version 3.1.1. https://mitsuba-renderer.org. 2022 (cit. on pp. 189, 202).

[Jam+19] Ondřej Jamriška, Šárka Sochorová, Ondřej Texler, Michal Lukáč, Jakub Fišer, Jingwan Lu, Eli Shechtman, and Daniel Sy`kora. “Stylizing Video by Example”. In: ACM Transactions on Graphics (TOG) 38.4 (2019), pp. 1–11 (cit. on p. 32).

[Jan+17] Michael Janner, Jiajun Wu, Tejas D Kulkarni, Ilker Yildirim, and Josh Tenenbaum. “Self-Supervised Intrinsic Image Decomposition”. In: Advances in Neural Information Processing Systems. 2017, pp. 5936–5946 (cit. on pp. 64, 67, 69, 78).

[JBH19] Miguel Jaques, Michael Burke, and Timothy Hospedales. “Physics-as-Inverse-Graphics: Unsupervised Physical Parameter Estimation from Video”. In: International Conference on Learning Representations (ICLR). 2019 (cit. on p. 158).

[JCJ09] Wojciech Jarosz, Nathan A. Carr, and Henrik Wann Jensen. “Importance Sampling Spherical Harmonics”. In: Computer Graphics Forum (Proceedings of Eurographics) 28.2 (2009), pp. 577–586 (cit. on p. 185).

[JBV16] Nikolay Jetchev, Urs Bergmann, and Roland Vollgraf. “Texture Synthesis with Spatial Generative Adversarial Networks”. In: arXiv preprint arXiv:1611.08207 (2016) (cit. on pp. 72, 74).

[Jia+21] Liming Jiang, Bo Dai, Wayne Wu, and Chen Change Loy. “Focal Frequency Loss for Image Reconstruction and Synthesis”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 13919–13929 (cit. on pp. 107, 128).

[Jin+22] Wenhua Jin, Beibei Wang, Miloš Hašan, Yu Guo, Steve Marschner, and Ling-Qi Yan. “Woven Fabric Capture from a Single Photo”. In: Proceedings of SIGGRAPH Asia 2022. 2022 (cit. on p. 125).

[Jin+19] Yongcheng Jing, Yezhou Yang, Zunlei Feng, Jingwen Ye, Yizhou Yu, and Mingli Song. “Neural Style Transfer: A review”. In: IEEE Transactions on Visualization and Computer Graphics (2019) (cit. on p. 32).

[JAF16] Justin Johnson, Alexandre Alahi, and Li Fei-Fei. “Perceptual Losses for Real-time Style Transfer and Super-resolution”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2016, pp. 694–711 (cit. on pp. 32, 74).

[Jos+22] Laurent Valentin Jospin, Hamid Laga, Farid Boussaid, Wray Buntine, and Mohammed Bennamoun. “Hands-on Bayesian Neural Networks—A Tutorial for Deep Learning Users”. In: IEEE Computational Intelligence Magazine 17.2 (2022), pp. 29–48 (cit. on p. 126).

[JC20] Eunjung Ju and Myung Geol Choi. “Estimating Cloth Simulation Parameters From a Static Drape Using Neural Networks”. In: IEEE Access 8 (2020), pp. 195113–195121 (cit. on pp. 156, 159).

[KBH03] Florian Kainz, Rod Bogart, and Drew Hess. “The OpenEXR Image File Format”. In: ACM SIGGRAPH Technical Sketches (2003) (cit. on p. 185).

[Kaj86] James T Kajiya. “The Rendering Equation”. In: Proceedings of the 13th annual conference on Computer graphics and interactive techniques. 1986, pp. 143–150 (cit. on pp. 4, 8).

[Kam+16] Christos Kampouris, Stefanos Zafeiriou, Abhijeet Ghosh, and Sotiris Malassiotis. “Fine-Grained Material Classification Using Micro-Geometry and Reflectance”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2016, pp. 778–792 (cit. on p. 40).

[KG13] Brian Karis and Epic Games. “Real Shading in Unreal Engine 4”. In: Proc. Physically Based Shading Theory Practice 4.3 (2013), p. 1 (cit. on p. 127).

[Kar+18] Tero Karras, Timo Aila, Samuli Laine, and Jaakko Lehtinen. “Progressive Growing of GANs for Improved Quality, Stability, and Variation”. In: International Conference on Learning Representations (ICLR). 2018 (cit. on pp. 74, 80).

[Kar+20a] Tero Karras, Miika Aittala, Janne Hellsten, Samuli Laine, Jaakko Lehtinen, and Timo Aila. “Training Generative Adversarial Networks with Limited Data”. In: arXiv preprint arXiv:2006.06676 (2020) (cit. on pp. 36, 109).

[Kar+21] Tero Karras, Miika Aittala, Samuli Laine, Erik Härkönen, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. “Alias-free Generative Adversarial Networks”. In: Advances in Neural Information Processing Systems. Vol. 34. 2021, pp. 852–863 (cit. on p. 186).

[KLA19] Tero Karras, Samuli Laine, and Timo Aila. “A Style-Based Generator Architecture for Generative Adversarial Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4401–4410 (cit. on p. 74).

[Kar+20b] Tero Karras, Samuli Laine, Miika Aittala, Janne Hellsten, Jaakko Lehtinen, and Timo Aila. “Analyzing and Improving the Image Quality of StyleGAN”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2020 (cit. on pp. 74, 85).

[Kas+15] Alexandre Kaspar, Boris Neubert, Dani Lischinski, Mark Pauly, and Johannes Kopf. “Self Tuning Texture Optimization”. In: Computer Graphics Forum. Vol. 34. 2. 2015, pp. 349–359 (cit. on p. 74).

[KZP19] Sergey Kastryulin, Dzhamil Zakirov, and Denis Prokopenko. PyTorch Image Quality: Metrics and Measure for Image Quality Assessment. Open-source software available at https://github.com/photosynthesis-team/piq. 2019 (cit. on pp. 182, 198).

[Kau17] Eric Kauderer-Abrams. “Quantifying Translation-invariance in Convolutional Neural Networks”. In: arXiv preprint arXiv:1801.01450 (2017) (cit. on p. 37).

[Kau+00] Jan Kautz, Pere-Pau Vázquez, Wolfgang Heidrich, and Hans-Peter Seidel. “Unified Approach to Prefiltered Environment Maps”. In: Proceedings of the Eurographics Workshop on Rendering Techniques 2000. 2000, pp. 185–196 (cit. on p. 185).

[Kaw80] Sueo Kawabata. “The Standardization and Analysis of Hand Evaluation”. In: The Textile Machinery Society of Japan (1980) (cit. on pp. 156, 158).

[KG17] Alex Kendall and Yarin Gal. “What Uncertainties Do We Need in Bayesian Deep Learning for Computer Vision?” In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 126).

[Khe+21] Ilyes Khemakhem, Ricardo Monti, Robert Leech, and Aapo Hyvarinen. “Causal Autoregressive Flows”. In: International Conference on Artificial Intelligence and Statistics. PMLR. 2021, pp. 3520–3528 (cit. on p. 191).

[KB15] Diederik P. Kingma and Jimmy Ba. “Adam: A Method for Stochastic Optimization”. In: International Conference on Learning Representations (ICLR). 2015 (cit. on pp. 39, 55, 67, 83, 94, 110, 142, 194).

[Kin+16] Durk P Kingma, Tim Salimans, Rafal Jozefowicz, Xi Chen, Ilya Sutskever, and Max Welling. “Improved Variational Inference with Inverse Autoregressive Flow”. In: Advances in Neural Information Processing Systems 29 (2016) (cit. on p. 191).

[KD18] Durk P. Kingma and Prafulla Dhariwal. “Glow: Generative Flow with Invertible 1x1 Convolutions”. In: Advances in Neural Information Processing Systems. Vol. 31. 2018 (cit. on pp. 186, 191).

[Kir+23] Alexander Kirillov, Eric Mintun, Nikhila Ravi, Hanzi Mao, Chloe Rolland, Laura Gustafson, Tete Xiao, Spencer Whitehead, Alexander C Berg, Wan-Yen Lo, et al. “Segment Anything”. In: arXiv preprint arXiv:2304.02643 (2023) (cit. on p. 212).

[KPB20] Ivan Kobyzev, Simon JD Prince, and Marcus A Brubaker. “Normalizing Flows: An Introduction and Review of Current Methods”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 43.11 (2020), pp. 3964–3979 (cit. on p. 191).

[Kol+20] Alexander Kolesnikov, Lucas Beyer, Xiaohua Zhai, Joan Puigcerver, Jessica Yung, Sylvain Gelly, and Neil Houlsby. “Big Transfer (bit): General Visual Representation Learning”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2020, pp. 491–507 (cit. on p. 159).

[KSL19] Simon Kornblith, Jonathon Shlens, and Quoc V Le. “Do Better Imagenet Models Transfer Better?” In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 2661–2671 (cit. on p. 159).

[Kou+03] Melissa L Koudelka, Sebastian Magda, Peter N Belhumeur, and David J Kriegman. “Acquisition, Compression, and Synthesis of Bidirectional Texture Functions”. In: 3rd International Workshop on Texture Analysis and Synthesis (Texture 2003). 2003, pp. 59–64 (cit. on p. 103).

[Kry+21] Michael C Krygier, Tyler LaBonte, Carianne Martinez, Chance Norris, Krish Sharma, Lincoln N Collins, Partha P Mukherjee, and Scott A Roberts. “Quantifying the Unknown Impact of Segmentation Uncertainty on Image-Based Simulations”. In: Nature Communications 12.1 (2021), pp. 1–11 (cit. on p. 126).

[KLG20] Sandra Kuijpers, Christiane Luible-Bär, and R. Hugh Gong. “The Measurement of Fabric Properties for Virtual Simulation—A Critical Review”. In: IEEE SA INDUSTRY CONNECTIONS (2020), pp. 1–43 (cit. on p. 158).

[KL51] Solomon Kullback and Richard A Leibler. “On Information and Sufficiency”. In: The Annals of Mathematical Statistics 22.1 (1951), pp. 79–86 (cit. on pp. 196, 197).

[Kum+22] Ananya Kumar, Aditi Raghunathan, Robbie Jones, Tengyu Ma, and Percy Liang. “Fine-tuning Can Distort Pretrained Features and Underperform Out-of-distribution”. In: arXiv preprint arXiv:2202.10054 (2022) (cit. on p. 159).

[Kum+19] Manoj Kumar, Mohammad Babaeizadeh, Dumitru Erhan, Chelsea Finn, Sergey Levine, Laurent Dinh, and Durk Kingma. “Videoflow: A Flow-based Generative Model for Video”. In: arXiv preprint arXiv:1903.01434 2.5 (2019), p. 3 (cit. on p. 186).

[Kur+19] Karol Kurach, Mario Lučić, Xiaohua Zhai, Marcin Michalski, and Sylvain Gelly. “A Large-Scale Study on Regularization and Normalization in GANs”. In: International Conference on Machine Learning (ICML). PMLR. 2019, pp. 3581–3590 (cit. on pp. 77, 94).

[Kur+22] Alexander Kurz, Katja Hauser, Hendrik Alexander Mehrtens, Eva Krieghoff-Henning, Achim Hekler, Jakob Nikolas Kather, Stefan Fröhling, Christof von Kalle, Titus Josef Brinker, et al. “Uncertainty Estimation in Medical Image Classification: Systematic Review”. In: JMIR Medical Informatics 10.8 (2022) (cit. on p. 126).

[Kuz+21] Alexandr Kuznetsov, Krishna Mullia, Zexiang Xu, Miloš Hašan, and Ravi Ramamoorthi. “NeuMIP: Multi-Resolution Neural Materials”. In: ACM Transactions on Graphics (TOG) 40.4 (2021), pp. 1–13 (cit. on pp. 18, 102, 103, 106, 107, 109–116, 119, 120, 187, 207, 259, 264).

[Kuz+22] Alexandr Kuznetsov, Xuezheng Wang, Krishna Mullia, Fujun Luan, Zexiang Xu, Milos Hasan, and Ravi Ramamoorthi. “Rendering Neural Materials on Curved Surfaces”. In: ACM SIGGRAPH 2022 Conference Proceedings. 2022, pp. 1–9 (cit. on pp. 102, 104, 116, 187).

[Kwa+05] Vivek Kwatra, Irfan Essa, Aaron Bobick, and Nipun Kwatra. “Texture Optimization for Example-Based Synthesis”. In: ACM Transactions on Graphics (TOG). Vol. 24. 3. ACM. 2005, pp. 795–802 (cit. on pp. 72, 74).

[Kwa+03] Vivek Kwatra, Arno Schödl, Irfan Essa, Greg Turk, and Aaron Bobick. “Graphcut Textures: Image and Video Synthesis Ysing Graph Cuts”. In: ACM Transactions on Graphics (TOG) 22.3 (2003), pp. 277–286 (cit. on p. 74).

[LGG19] Manuel Lagunas, Elena Garces, and Diego Gutierrez. “Learning Icons Appearance Similarity”. In: Multimedia Tools and Applications 78.8 (2019), pp. 10733–10751 (cit. on p. 160).

[Lag+19] Manuel Lagunas, Sandra Malpica, Ana Serrano, Elena Garces, Diego Gutierrez, and Belen Masia. “A Similarity Measure for Material Appearance”. In: ACM Transactions on Graphics (TOG) 38.4 (2019), pp. 1–12 (cit. on pp. 157, 160, 180).

[LKA13] Samuli Laine, Tero Karras, and Timo Aila. “Megakernels Considered Harmful: Wavefront Path Tracing on GPUs”. In: Proceedings of the 5th High-Performance Graphics Conference. 2013, pp. 137–143 (cit. on pp. 109, 195).

[LPB17] Balaji Lakshminarayanan, Alexander Pritzel, and Charles Blundell. “Simple and Scalable Predictive Uncertainty Estimation Using Deep Ensembles”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on pp. 126, 130).

[LS98] Greg Ward Larson and Rob Shakespeare. Rendering with Radiance: The Art and Science of Lighting Visualization. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 1998 (cit. on p. 185).

[Lav+21] Guillaume Lavoué, Nicolas Bonneel, Jean-Philippe Farrugia, and Cyril Soler. “Perceptual Quality of BRDF Approximations: Dataset and Metrics”. In: Computer Graphics Forum. Vol. 40. 2. Wiley Online Library. 2021, pp. 327–338 (cit. on pp. 130, 131, 193).

[Led+17] Christian Ledig, Lucas Theis, Ferenc Huszár, et al. “Photo-Realistic Single Image Super-Resolution Using a Generative Adversarial Network”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) 2017-Janua (Sept. 2017), pp. 105–114 (cit. on pp. 79, 94).

[LH06] Sylvain Lefebvre and Hugues Hoppe. “Appearance-Space Texture Synthesis”. In: ACM Transactions on Graphics (TOG) 25.3 (2006), pp. 541–548 (cit. on pp. 34, 104).

[Lei+21] Zhao Lei, Yi Zeng, Peng Liu, and Xiaohui Su. “Active Deep Learning for Hyperspectral Image Classification with Uncertainty Learning”. In: IEEE Geoscience and Remote Sensing Letters 19 (2021) (cit. on p. 126).

[LM01] Thomas Leung and Jitendra Malik. “Representing and Recognizing the Visual Appearance of Materials Using Three-Dimensional Textons”. In: International Journal of Computer Vision 43.1 (2001), pp. 29–44 (cit. on pp. 33, 39).

[LLB22] Wei-Hong Li, Xialei Liu, and Hakan Bilen. “Cross-domain Few-shot Learning with Task-specific Adapters”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition. 2022, pp. 7161–7170 (cit. on p. 159).

[Li+17a] Xiao Li, Yue Dong, Pieter Peers, and Xin Tong. “Modeling Surface Appearance from a Single Photograph Using Self-Augmented Convolutional Neural Networks”. In: ACM Transactions on Graphics (TOG) 36.4 (2017), pp. 1–11 (cit. on pp. 34, 124).

[Li+19] Xiaoyu Li, Bo Zhang, Pedro V Sander, and Jing Liao. “Blind Geometric Distortion Correction on Images through Deep Learning”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4855–4864 (cit. on p. 93).

[Li+22] Yifei Li, Tao Du, Kui Wu, Jie Xu, and Wojciech Matusik. “DiffCloth: Differentiable Cloth Simulation with Dry Frictional Contact”. In: ACM Transactions on Graphics (TOG) (2022) (cit. on p. 158).

[Li+17b] Yijun Li, Chen Fang, Jimei Yang, Zhaowen Wang, Xin Lu, and Ming-Hsuan Yang. “Universal Style Transfer via Feature Transforms”. In: Advances in Neural Information Processing Systems. 2017, pp. 386–396 (cit. on p. 32).

[LAA08] Yuanzhen Li, Edward Adelson, and Aseem Agarwala. “ScribbleBoost: Adding Classification to Edge-Aware Interpolation of Local Image and Video Adjustments”. In: Computer Graphics Forum. Vol. 27. 4. Wiley Online Library. 2008, pp. 1255–1264 (cit. on p. 33).

[LS18] Zhengqi Li and Noah Snavely. “Learning Intrinsic Image Decomposition From Watching the World”. In: Proceedings of the IEEE Conference on Computer Vision and Pattern Recognition (CVPR). June 2018 (cit. on p. 78).

[Li+20] Zhengqin Li, Mohammad Shafiei, Ravi Ramamoorthi, Kalyan Sunkavalli, and Man-mohan Chandraker. “Inverse Rendering for Complex Indoor Scenes: Shape, Spatially-Varying Lighting and SVBRDF From a Single Image”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 2475–2484 (cit. on pp. 72, 74, 87–91, 95, 99, 114, 203).

[LSC18] Zhengqin Li, Kalyan Sunkavalli, and Manmohan Chandraker. “Materials for Masses: SVBRDF Acquisition with a Single Mobile Phone Image”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 72–87 (cit. on pp. 72, 125).

[Li+18] Zhengqin Li, Zexiang Xu, Ravi Ramamoorthi, Kalyan Sunkavalli, and Manmohan Chandraker. “Learning to Reconstruct Shape and Spatially-varying Reflectance from a Single Image”. In: ACM Transactions on Graphics (TOG) 37.6 (2018), pp. 1–11 (cit. on pp. 125, 129).

[LLK19] Junbang Liang, Ming Lin, and Vladlen Koltun. “Differentiable Cloth Simulation for Inverse Problems”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on p. 158).

[Lia+17] Jing Liao, Yuan Yao, Lu Yuan, Gang Hua, and Sing Bing Kang. “Visual Attribute Transfer through Deep Image Analogy”. In: ACM Transactions on Graphics (TOG) 36.4 (2017), pp. 1–15 (cit. on pp. 31, 32, 46, 47, 56, 57, 60).

[Lic+21] Daniel Lichy, Jiaye Wu, Soumyadip Sengupta, and David W Jacobs. “Shape and Material Capture at Home”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 6123–6133 (cit. on p. 129).

[LPG19] Yiming Lin, Pieter Peers, and Abhijeet Ghosh. “On-Site Example-Based Material Appearance Acquisition”. In: Computer Graphics Forum. Vol. 38. 4. Wiley Online Library. 2019, pp. 15–25 (cit. on pp. 34, 105).

[Liu+20] Guilin Liu, Rohan Taori, Ting-Chun Wang, Zhiding Yu, Shiqiu Liu, Fitsum A Reda, Karan Sapra, Andrew Tao, and Bryan Catanzaro. “Transposer: Universal Texture Synthesis Using Feature Maps as Transposed Convolution Filter”. In: arXiv preprint arXiv:2007.07243 (2020) (cit. on pp. 72, 74, 75, 92).

[Liu+19a] Liyuan Liu, Haoming Jiang, Pengcheng He, Weizhu Chen, Xiaodong Liu, Jianfeng Gao, and Jiawei Han. “On the Variance of the Adaptive Learning Rate and Beyond”. In: arXiv preprint arXiv:1908.03265 (2019) (cit. on p. 110).

[Liu+19b] Ming-Yu Liu, Xun Huang, Arun Mallya, Tero Karras, Timo Aila, Jaakko Lehtinen, and Jan Kautz. “Few-shot Unsupervised Image-to-Image Translation”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2019, pp. 10551–10560 (cit. on p. 37).

[Liu+21] Steven Liu, Xiuming Zhang, Zhoutong Zhang, Richard Zhang, Jun-Yan Zhu, and Bryan Russell. “Editing conditional radiance fields”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 5773–5783 (cit. on p. 116).

[Liu+22a] Yuan Liu, Sida Peng, Lingjie Liu, Qianqian Wang, Peng Wang, Christian Theobalt, Xiaowei Zhou, and Wenping Wang. “Neural Rays for Occlusion-aware Image-based Rendering”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 7824–7833 (cit. on p. 203).

[Liu+22b] Zhuang Liu, Hanzi Mao, Chao-Yuan Wu, Christoph Feichtenhofer, Trevor Darrell, and Saining Xie. “A Convnet for the 2020s”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 11976–11986 (cit. on pp. 108, 109, 119).

[LSD15] Jonathan Long, Evan Shelhamer, and Trevor Darrell. “Fully Convolutional Networks for Semantic Segmentation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2015, pp. 3431–3440 (cit. on pp. 83, 94).

[LH17] Ilya Loshchilov and Frank Hutter. “Decoupled Weight Decay Regularization”. In: arXiv preprint arXiv:1711.05101 (2017) (cit. on p. 176).

[Lua+21] Fujun Luan, Shuang Zhao, Kavita Bala, and Zhao Dong. “Unified Shape and SVBRDF Recovery Using Differentiable Monte Carlo Rendering”. In: Computer Graphics Forum. Vol. 40. 4. Wiley Online Library. 2021, pp. 101–113 (cit. on pp. 125, 129).

[Lug+20] Andreas Lugmayr, Martin Danelljan, Luc Van Gool, and Radu Timofte. “Srflow: Learning the Super-resolution Space with Normalizing Flow”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2020, pp. 715–732 (cit. on p. 186).

[LM08] Christiane Luible and Nadia Magnenat-Thalmann. “The Simulation of Cloth Using Accurate Physical Parameters”. In: Proceedings of the Tenth IASTED International Conference on Computer Graphics and Imaging (CGIM ’08). 2008 (cit. on p. 158).

[LKS15] Zhaoliang Lun, Evangelos Kalogerakis, and Alla Sheffer. “Elements of Style: Learning Perceptual Shape Style Similarity”. In: ACM Transactions on graphics (TOG) 34.4 (2015), pp. 1–14 (cit. on p. 160).

[MHN13] Andrew L Maas, Awni Y Hannun, and Andrew Y Ng. “Rectifier Nonlinearities Improve Neural Network Acoustic Models”. In: International Conference on Machine Learning (ICML). Vol. 30. 1. Citeseer. 2013, p. 3 (cit. on pp. 77, 94, 176).

[Mag+07] Nadia Magnenat-Thalmann, Christiane Luible, Pascal Volino, and Etienne Lyard. “From Measured Fabric to the Simulation of Cloth”. In: 2007 10th IEEE International Conference on Computer-Aided Design and Computer Graphics. IEEE. 2007, pp. 7–18 (cit. on p. 158).

[MR10] Sébastien Marcel and Yann Rodriguez. “Torchvision the Machine-vision Package of Torch”. In: Proceedings of the 18th ACM International Conference on Multimedia. 2010, pp. 1485–1488 (cit. on pp. 67, 84, 110, 142, 194).

[Mar+20] Morteza Mardani, Guilin Liu, Aysegul Dundar, Shiqiu Liu, Andrew Tao, and Bryan Catanzaro. “Neural FFTs for Universal Texture Image Synthesis”. In: Advances in Neural Information Processing Systems 33 (2020) (cit. on pp. 72, 74, 75, 85, 107, 128).

[MFS08] Bruce A Maxwell, Richard M Friedhoff, and Casey A Smith. “A biilluminant dichromatic reflection model for understanding images”. In: 2008 IEEE Conference on Computer Vision and Pattern Recognition. IEEE. 2008, pp. 1–8 (cit. on p. 10).

[Maz+19] Ilya Mazlov, Sebastian Merzbach, Elena Trunz, and Reinhard Klein. “Neural Appearance Synthesis and Transfer”. In: (2019), pp. 35–39 (cit. on p. 31).

[MHM18] Leland McInnes, John Healy, and James Melville. “Umap: Uniform Manifold Approximation and Projection for Dimension Reduction”. In: arXiv preprint arXiv:1802.03426 (2018) (cit. on p. 133).

[Meh+21] Ishit Mehta, Michaël Gharbi, Connelly Barnes, Eli Shechtman, Ravi Ramamoorthi, and Manmohan Chandraker. “Modulated Periodic Activations for Generalizable Local Functional Representations”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 14214–14223 (cit. on pp. 109, 186, 203).

[MR21] Sachin Mehta and Mohammad Rastegari. “MobileViT: Light-Weight, General-Purpose, and Mobile-Friendly Vision Transformer”. In: International Conference on Learning Representations (ICLR). 2021 (cit. on pp. 124, 127, 132, 141, 144).

[Mel+12] Francho Melendez, Mashhuda Glencross, Jon Starck, and Gregory J Ward. “Transfer of Albedo and Local Depth Variation to Photo-Textures”. In: Proceedings of the 9th European Conference on Visual Media Production. 2012, pp. 40–48 (cit. on pp. 31, 34).

[Mel+21] Joe Mellor, Jack Turner, Amos Storkey, and Elliot J Crowley. “Neural Architecture Search Without Training”. In: International Conference on Machine Learning (ICML). PMLR. 2021, pp. 7588–7598 (cit. on p. 92).

[Mer+19] Sebastian Merzbach, Max Hermann, Martin Rump, and Reinhard Klein. “Learned Fitting of Spatially Varying BRDFs”. In: Computer Graphics Forum. Vol. 38. 4. Wiley Online Library. 2019, pp. 193–205 (cit. on p. 30).

[MWK17] Sebastian Merzbach, Michael Weinmann, and Reinhard Klein. “High-quality Multi-Spectral Reflectance Acquisition with X-rite TAC7”. In: Proceedings of the Workshop on Material Appearance Modeling. 2017, pp. 11–16 (cit. on pp. 35, 39).

[MNG17] Lars Mescheder, Sebastian Nowozin, and Andreas Geiger. “Adversarial Variational Bayes: Unifying Variational Autoencoders and Generative Adversarial Networks”. In: International Conference on Machine Learning (ICML). PMLR. 2017, pp. 2391–2400 (cit. on p. 126).

[Mic+18] Paulius Micikevicius, Sharan Narang, Jonah Alben, Gregory Diamos, Erich Elsen, David Garcia, Boris Ginsburg, Michael Houston, Oleksii Kuchaiev, Ganesh Venkatesh, et al. “Mixed Precision Training”. In: International Conference on Learning Representations (ICLR). 2018 (cit. on pp. 56, 67, 84, 94, 110, 142, 176, 194).

[Mig+12] Eder Miguel, Derek Bradley, Bernhard Thomaszewski, Bernd Bickel, Wojciech Matusik, Miguel A Otaduy, and Steve Marschner. “Data-driven Estimation of Cloth Simulation Models”. In: Computer Graphics Forum. Vol. 31. 2pt2. Wiley Online Library. 2012, pp. 519–528 (cit. on p. 158).

[Mil+20] Ben Mildenhall, Pratul P. Srinivasan, Matthew Tancik, Jonathan T. Barron, Ravi Ramamoorthi, and Ren Ng. “NeRF: Representing Scenes as Neural Radiance Fields for View Synthesis”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2020 (cit. on pp. 186, 187, 207, 212, 262, 264).

[Min95] Pier Giorgio Minazio. “FAST–Fabric Assurance by Simple Testing”. In: International Journal of Clothing Science and Technology (1995) (cit. on pp. 156, 158).

[Miy+18] Takeru Miyato, Toshiki Kataoka, Masanori Koyama, and Yuichi Yoshida. “Spectral Normalization for Generative Adversarial Networks”. In: International Conference on Learning Representations. 2018 (cit. on pp. 141, 145).

[Mor+20] Alexander Mordvintsev, Ettore Randazzo, Eyvind Niklasson, and Michael Levin. “Growing Neural Cellular Automata”. In: Distill (2020). https://distill.pub/2020/growing-ca (cit. on pp. 15, 75, 257).

[Mor15] Andrew Morgan. The True Cost. 2015 (cit. on p. 2).

[Mor+17] Joep Moritz, Stuart James, Tom S.F. Haines, Tobias Ritschel, and Tim Weyrich. “Texture Stationarization: Turning Photos into Tileable Textures”. In: Eurographics Symposium on Geometry Processing. Vol. 36. 2. 2017, pp. 177–188 (cit. on pp. 15, 72–74, 88–90, 99, 114, 257).

[MMK03] Gero Müller, Jan Meseth, and Reinhard Klein. “Compression and Real-Time Rendering of Measured BTFs Using Local PCA.” In: VMV. 2003, pp. 271–279 (cit. on p. 103).

[Mül+22] Thomas Müller, Alex Evans, Christoph Schied, and Alexander Keller. “Instant Neural Graphics Primitives with a Multiresolution Hash Encoding”. In: ACM Transactions on Graphics (TOG) 41.4 (2022), 102:1–102:15 (cit. on pp. 203, 212).

[Mül+19] Thomas Müller, Brian McWilliams, Fabrice Rousselle, Markus Gross, and Jan Novák. “Neural Importance Sampling”. In: ACM Transactions on Graphics (TOG) 38.5 (2019), pp. 1–19 (cit. on pp. 116, 186, 191, 196, 197).

[Mül+20] Thomas Müller, Fabrice Rousselle, Alexander Keller, and Jan Novák. “Neural Control Variates”. In: ACM Transactions on Graphics (TOG) 39.6 (2020), pp. 1–19 (cit. on pp. 186, 191).

[Mun+22] Jacob Munkberg, Jon Hasselgren, Tianchang Shen, Jun Gao, Wenzheng Chen, Alex Evans, Thomas Müller, and Sanja Fidler. “Extracting Triangular 3D Models, Materials, and Lighting From Images”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 8280–8290 (cit. on pp. 127, 129).

[Mur+20] J Krishna Murthy, Miles Macklin, Florian Golemo, Vikram Voleti, Linda Petrini, Martin Weiss, Breandan Considine, Jérôme Parent-Lévesque, Kevin Xie, Kenny Erleben, et al. “GradSim: Differentiable simulation for System Identification and Visuomotor Control”. In: International Conference on Learning Representations (ICLR). 2020 (cit. on p. 158).

[Naf+16] Hossein Ziaei Nafchi, Atena Shahkolaei, Rachid Hedjam, and Mohamed Cheriet. “Mean Deviation Similarity Index: Efficient and Reliable Full-reference Image Quality Evaluator”. In: IEEE Access 4 (2016), pp. 5579–5590 (cit. on p. 160).

[Nag+15] Koki Nagano, Graham Fyffe, Oleg Alexander, Jernej Barbic, Hao Li, Abhijeet Ghosh, and Paul E Debevec. “Skin Microstructure Deformation with Displacement Map Convolution.” In: ACM Transactions on Graphics (TOG) 34.4 (2015), pp. 109–1 (cit. on p. 105).

[NH10] Vinod Nair and Geoffrey E Hinton. “Rectified Linear Units Improve Restricted Boltzmann Machines”. In: International Conference on Machine Learning (ICML). 2010 (cit. on pp. 54, 109, 111, 120).

[Nam+16] Giljoo Nam, Joo Ho Lee, Hongzhi Wu, Diego Gutierrez, and Min H Kim. “Simultaneous Acquisition of Microscale Reflectance and Normals”. In: ACM Transactions on Graphics (TOG) 35.6 (2016), pp. 1–11 (cit. on pp. 35, 39).

[NK18] Hyeonseob Nam and Hyo Eun Kim. “Batch-Instance Normalization for Adaptively Style-Invariant Neural Networks”. In: Advances in Neural Information Processing Systems. Vol. 2018-Decem. 2018, pp. 2558–2567 (cit. on p. 77).

[NSO12] Rahul Narain, Armin Samii, and James F O’brien. “Adaptive Anisotropic Remeshing for Cloth Simulation”. In: ACM Transactions on Graphics (TOG) 31.6 (2012), pp. 1–10 (cit. on p. 162).

[ND06] Addy Ngan and Frédo Durand. “Statistical Acquisition of Texture Appearance”. In: Proceedings of the 17th Eurographics Conference on Rendering Techniques. 2006, pp. 31–40 (cit. on p. 63).

[NJR15] Jannik Boll Nielsen, Henrik Wann Jensen, and Ravi Ramamoorthi. “On Optimal, Minimal BRDF Sampling for Reflectance Acquisition”. In: ACM Transactions on Graphics (TOG) 34.6 (2015), pp. 1–11 (cit. on pp. 42, 116, 130).

[Nik+21] Eyvind Niklasson, Alexander Mordvintsev, Ettore Randazzo, and Michael Levin. “Self-Organising Textures”. In: Distill (2021). https://distill.pub/selforg/2021/textures (cit. on pp. 72, 74, 75, 89–91, 93, 95, 99).

[Nim+19] Merlin Nimier-David, Delio Vicini, Tizian Zeltner, and Wenzel Jakob. “Mitsuba 2: A Retargetable Forward and Inverse Renderer”. In: ACM Transactions on Graphics (TOG) 38.6 (2019), pp. 1–17 (cit. on pp. 6, 109, 202).

[ODO16] Augustus Odena, Vincent Dumoulin, and Chris Olah. “Deconvolution and Checker-board Artifacts”. In: Distill (2016) (cit. on p. 109).

[Ouy+21] Yaobin Ouyang, Shiqiu Liu, Markus Kettunen, Matt Pharr, and Jacopo Pantaleoni. “ReSTIR GI: Path Resampling for Real-time Path Tracing”. In: Computer Graphics Forum. Vol. 40. 8. 2021, pp. 17–29 (cit. on p. 184).

[Pap+21] George Papamakarios, Eric Nalisnick, Danilo Jimenez Rezende, Shakir Mohamed, and Balaji Lakshminarayanan. “Normalizing Flows for Probabilistic Modeling and Inference”. In: The Journal of Machine Learning Research 22.1 (2021), pp. 2617–2680 (cit. on p. 191).

[PPM17] George Papamakarios, Theo Pavlakou, and Iain Murray. “Masked Autoregressive Flow for Density Estimation”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 191).

[PSM19] George Papamakarios, David Sterratt, and Iain Murray. “Sequential Neural Likelihood: Fast Likelihood-free Inference with Autoregressive Flows”. In: The 22nd International Conference on Artificial Intelligence and Statistics. PMLR. 2019, pp. 837–848 (cit. on p. 191).

[Par+20] Taesung Park, Alexei A Efros, Richard Zhang, and Jun-Yan Zhu. “Contrastive Learning for Unpaired Image-to-Image Translation”. In: arXiv preprint arXiv:2007.15651 (2020) (cit. on pp. 37, 210).

[Pas+19] Adam Paszke, Sam Gross, Francisco Massa, Adam Lerer, James Bradbury, Gregory Chanan, Trevor Killeen, Zeming Lin, Natalia Gimelshein, Luca Antiga, et al. “Pytorch: An Imperative Style, High-performance Deep Learning Library”. In: Advances in Neural Information Processing Systems. Vol. 32. 2019 (cit. on pp. 39, 55, 67, 84, 94, 110, 142, 176, 194).

[Pha+20] Minh Quang Pham, Josep-Maria Crego, François Yvon, and Jean Senellart. “A study of Residual Adapters for Multi-domain Neural Machine Translation”. In: Conference on Machine Translation. 2020 (cit. on p. 159).

[Pha19] Matt Pharr. Visualizing Warping Strategies for Sampling Environment Map Lights. 2019 (cit. on pp. 189, 199, 200).

[PJH20] Matt Pharr, Wenzel Jakob, and Greg Humphreys. Implementation of the forthcoming 4th edition of Physically Based Rendering: From Theory to Implementation. 2020 (cit. on pp. 3, 109, 189, 200).

[PJH16] Matt Pharr, Wenzel Jakob, and Greg Humphreys. Physically Based Rendering: From Theory to Implementation. 3rd. San Francisco, CA, USA: Morgan Kaufmann Publishers Inc., 2016 (cit. on pp. 185, 189, 200, 201).

[PS00] Javier Portilla and Eero P Simoncelli. “A Parametric Texture Model Based on Joint Statistics of Complex Wavelet Coefficients”. In: International Journal of Computer Vision 40.1 (2000), pp. 49–70 (cit. on p. 74).

[Pow13] Jess Power. “Fabric Objective Measurements for Commercial 3D Virtual Garment Simulation”. In: International Journal of Clothing Science and Technology (2013) (cit. on p. 158).

[Pra+18] Ekta Prashnani, Hong Cai, Yasamin Mostofi, and Pradeep Sen. “Pieapp: Perceptual Image-error Assessment Through Pairwise Preference”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 1808–1817 (cit. on p. 160).

[Raa+18] Lara Raad, Axel Davy, Agnès Desolneux, and Jean-Michel Morel. “A Survey of Exemplar-Based Texture Synthesis”. In: Annals of Mathematical Sciences and Applications 3.1 (2018), pp. 89–148 (cit. on p. 73).

[Rad+21] Alec Radford, Jong Wook Kim, Chris Hallacy, Aditya Ramesh, Gabriel Goh, Sandhini Agarwal, Girish Sastry, Amanda Askell, Pamela Mishkin, Jack Clark, et al. “Learning Transferable Visual Models from Natural Language Supervision”. In: International Conference on Machine Learning (ICML). PMLR. 2021, pp. 8748–8763 (cit. on p. 159).

[Rag+19] Maithra Raghu, Chiyuan Zhang, Jon Kleinberg, and Samy Bengio. “Transfusion: Under-standing Transfer Learning for Medical Imaging”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on p. 159).

[Rai+22] Gilles Rainer, Adrien Bousseau, Tobias Ritschel, and George Drettakis. “Neural Pre-computed Radiance Transfer”. In: Computer Graphics Forum. Vol. 41. 2. Wiley Online Library. 2022, pp. 365–378 (cit. on pp. 105, 116, 203).

[Rai+20] Gilles Rainer, Abhijeet Ghosh, Wenzel Jakob, and Tim Weyrich. “Unified Neural Encod-ing of BTFs”. In: Computer Graphics Forum. Vol. 39. 2. The Eurographics Association. 2020, pp. 1–13 (cit. on pp. 18, 33, 102, 103, 109, 112, 113, 116, 259).

[Rai+19] Gilles Rainer, Wenzel Jakob, Abhijeet Ghosh, and Tim Weyrich. “Neural BTF Compres-sion and Interpolation”. In: Computer Graphics Forum. Vol. 38. 2. Wiley Online Library. 2019, pp. 235–244 (cit. on pp. 18, 33, 39, 40, 102, 103, 106, 109, 112, 113, 259).

[Ram+19] Prajit Ramachandran, Niki Parmar, Ashish Vaswani, Irwan Bello, Anselm Levskaya, and Jon Shlens. “Stand-alone Self-attention in Vision Models”. In: Advances in Neural Information Processing Systems 32 (2019) (cit. on p. 164).

[RH01a] Ravi Ramamoorthi and Pat Hanrahan. “An Efficient Representation for Irradiance Environment Maps”. In: Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques. SIGGRAPH ’01. New York, NY, USA: Association for Computing Machinery, 2001, pp. 497–500 (cit. on p. 13).

[RH01b] Ravi Ramamoorthi and Pat Hanrahan. “An Efficient Representation for Irradiance Environment Maps”. In: Proceedings of the 28th Annual Conference on Computer Graphics and Interactive Techniques. 2001, pp. 497–500 (cit. on pp. 196, 199, 202).

[Ras+20] Abdullah Haroon Rasheed, Victor Romero, Florence Bertails-Descoubes, Stefanie Wuhrer, Jean-Sébastien Franco, and Arnaud Lazarus. “Learning to Measure the Static Friction Coefficient in Cloth Contact”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 9912–9921 (cit. on p. 159).

[RD17] Ajita Rattani and Reza Derakhshani. “On Fine-tuning Convolutional Neural Networks for Smartphone Based Ocular Recognition”. In: 2017 IEEE international joint conference on biometrics (IJCB). IEEE. 2017, pp. 762–767 (cit. on p. 159).

[RBV18] Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. “Efficient Parametrization of Multi-domain Deep Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 8119–8127 (cit. on p. 159).

[RBV17] Sylvestre-Alvise Rebuffi, Hakan Bilen, and Andrea Vedaldi. “Learning Multiple Visual Domains with Residual Adapters”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 159).

[RJ19] A Sai Bharadwaj Reddy and D Sujitha Juliet. “Transfer Learning with ResNet-50 for Malaria Cell-image Classification”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). IEEE. 2019, pp. 0945–0949 (cit. on p. 159).

[Ree+15] Scott E Reed, Yi Zhang, Yuting Zhang, and Honglak Lee. “Deep Visual Analogy-Making”. In: Advances in Neural Information Processing Systems. 2015, pp. 1252–1260 (cit. on p. 32).

[Rei+14] Klein Reinhard, Rump Martin, Weinmann Michael, Sarlette Ralf, and Schwartz Christopher. “Design and Implementation of Practical Bidirectional Texture Function Measurement Devices focusing on the Developments at the University of Bonn”. In: Sensors 14.5 (2014) (cit. on pp. 109, 114, 117).

[Rei+18] Rafael Reisenhofer, Sebastian Bosse, Gitta Kutyniok, and Thomas Wiegand. “A Haar Wavelet-based Perceptual Similarity Index for Image Quality Assessment”. In: Signal Processing: Image Communication 61 (2018), pp. 33–43 (cit. on p. 160).

[RSS16] Nathalie Remy, Eveline Speelman, and Steven Swartz. Style that’s Sustainable: A New Fast-Fashion Formula. 2016 (cit. on p. 2).

[Ren+17] Jian Ren, Xiaohui Shen, Zhe Lin, Radomir Mech, and David J Foran. “Personalized Image Aesthetics”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2017, pp. 638–647 (cit. on p. 74).

[Rez+20] Danilo Jimenez Rezende, George Papamakarios, Sébastien Racaniere, Michael Albergo, Gurtej Kanwar, Phiala Shanahan, and Kyle Cranmer. “Normalizing Flows on Tori and Spheres”. In: International Conference on Machine Learning (ICML). PMLR. 2020, pp. 8083–8092 (cit. on p. 189).

[Al-+16] Rami Al-Rfou, Guillaume Alain, Amjad Almahairi, Christof Angermueller, Dzmitry Bahdanau, Nicolas Ballas, Frédéric Bastien, Justin Bayer, Anatoly Belikov, Alexander Belopolsky, et al. “Theano: A Python Framework for Fast Computation of Mathematical Expressions”. In: arXiv e-prints (2016), arXiv–1605 (cit. on p. 95).

[Rho+22] Daniel Rho, Junwoo Cho, Jong Hwan Ko, and Eunbyung Park. “Neural Residual Flow Fields for Efficient Video Representations”. In: Proceedings of the Asian Conference on Computer Vision. 2022, pp. 3447–3463 (cit. on p. 187).

[Rib+20] Edgar Riba, Dmytro Mishkin, Daniel Ponsa, Ethan Rublee, and Gary Bradski. “Kornia: An Open Source Differentiable Computer Vision Library for Pytorch”. In: Proceedings of the IEEE/CVF Winter Conference on Applications of Computer Vision. 2020, pp. 3674–3683 (cit. on pp. 67, 111, 142, 176).

[Ric+16] Stephan R. Richter, Vibhav Vineet, Stefan Roth, and Vladlen Koltun. “Playing for Data: Ground Truth from Computer Games”. In: European Conference on Computer Vision (ECCV). Ed. by Bastian Leibe, Jiri Matas, Nicu Sebe, and Max Welling. Vol. 9906. LNCS. Springer International Publishing, 2016, pp. 102–118 (cit. on p. 212).

[RPG16] Jérémy Riviere, Pieter Peers, and Abhijeet Ghosh. “Mobile Surface Reflectometry”. In: Computer Graphics Forum. Vol. 35. 1. Wiley Online Library. 2016, pp. 191–202 (cit. on pp. 31, 34).

[Rod+23a] Carlos Rodriguez-Pardo, Henar Dominguez-Elvira, David Pascual-Hernandez, and Elena Garces. “UMat: Uncertainty-Aware Single Image High Resolution Material Capture”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2023) (cit. on pp. 24, 62, 76, 102, 107, 108, 110, 116, 193, 268).

[RG21] Carlos Rodriguez-Pardo and Elena Garces. “Neural Photometry-guided Visual Attribute Transfer”. In: IEEE Transactions on Visualization and Computer Graphics (2021) (cit. on pp. 23, 25, 63–70, 102, 105, 107, 108, 110, 114, 125, 129, 131, 142, 193, 268).

[RG22] Carlos Rodriguez-Pardo and Elena Garces. “SeamlessGAN: Self-Supervised Synthesis of Tileable Texture Maps”. In: IEEE Transactions on Visualization and Computer Graphics (2022) (cit. on pp. 23, 25, 64, 67, 69, 104, 107, 108, 110, 114, 127, 128, 141, 268).

[Rod+23b] Carlos Rodriguez-Pardo, Melania Prieto-Martin, Dan Casas, and Elena Garces. “How Will It Drape Like? Capturing Fabric Mechanics from Depth Images”. In: Computer Graphics Forum (Proc. of Eurographics) (2023) (cit. on pp. 18, 23, 212, 268).

[Rod+19] Carlos Rodriguez-Pardo, Sergio Suja, David Pascual, Jorge Lopez-Moreno, and Elena Garces. “Automatic Extraction and Synthesis of Regular Repeatable Patterns”. In: Computers & Graphics 83 (2019), pp. 33–41 (cit. on pp. 15, 34, 72, 75, 79, 83, 87, 89–91, 93, 95, 99, 114, 257).

[RB19] Carlos Rodrıguez-Pardo and Hakan Bilen. “Personalised Aesthetics with Residual Adapters”. In: Iberian Conference on Pattern Recognition and Image Analysis. Springer. 2019, pp. 508–520 (cit. on pp. 74, 90, 159).

[Rom+22] Robin Rombach, Andreas Blattmann, Dominik Lorenz, Patrick Esser, and Björn Ommer. “High-Resolution Image Synthesis with Latent Diffusion Models”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 10684–10695 (cit. on pp. 127, 209, 210).

[RFB15] Olaf Ronneberger, Philipp Fischer, and Thomas Brox. “U-Net: Convolutional Networks for Biomedical Image Segmentation”. In: International Conference on Medical image computing and computer-assisted intervention. Springer. 2015, pp. 234–241 (cit. on pp. 37–39, 54, 55, 64, 66, 67, 69, 78, 108, 111, 124, 127, 132, 141).

[RSK13] Roland Ruiters, Christopher Schwartz, and Reinhard Klein. “Example-based Interpolation and Synthesis of Bidirectional Texture Functions”. In: Computer Graphics Forum. Vol. 32. 2pt3. Wiley Online Library. 2013, pp. 361–370 (cit. on p. 63).

[RSK10] Martin Rump, Ralf Sarlette, and Reinhard Klein. “Groundtruth Data for Multispectral Bidirectional Texture Functions”. In: Conference on Colour in Graphics, Imaging, and Vision. Vol. 2010. 1. Society for Imaging Science and Technology. 2010, pp. 326–331 (cit. on p. 116).

[Run+20] Tom FH Runia, Kirill Gavrilyuk, Cees GM Snoek, and Arnold WM Smeulders. “Cloth in the Wind: A Case Study of Physical Measurement Through Simulation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 10498–10507 (cit. on p. 158).

[Rus+04] Daniel B Russakoff, Carlo Tomasi, Torsten Rohlfing, and Calvin R Maurer. “Image Similarity Using Mutual Information of Regions”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2004, pp. 596–607 (cit. on p. 132).

[Sah+22] Chitwan Saharia, William Chan, Huiwen Chang, Chris Lee, Jonathan Ho, Tim Salimans, David Fleet, and Mohammad Norouzi. “Palette: Image-to-Image Diffusion Models”. In: ACM SIGGRAPH 2022 Conference Proceedings. 2022, pp. 1–10 (cit. on pp. 127, 209).

[San+19] Veit Sandfort, Ke Yan, Perry J Pickhardt, and Ronald M Summers. “Data Augmentation Using Generative Adversarial Networks (CycleGAN) to Improve Generalizability in CT Segmentation Tasks”. In: Scientific reports 9.1 (2019), pp. 1–9 (cit. on p. 45).

[SC20] Shen Sang and Manmohan Chandraker. “Single-Shot Neural Relighting and SVBRDF Estimation”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2020, pp. 85–101 (cit. on p. 125).

[SOC22] Igor Santesteban, Miguel A Otaduy, and Dan Casas. “SNUG: Self-Supervised Neural Dynamic Garments”. In: IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR) (2022) (cit. on pp. 17, 259).

[SSK03] Mirko Sattler, Ralf Sarlette, and Reinhard Klein. “Efficient and Realistic Visualization of Cloth”. In: Rendering techniques. 2003, pp. 167–178 (cit. on p. 109).

[SSK20] Edgar Schonfeld, Bernt Schiele, and Anna Khoreva. “A U-Net Based Discriminator for Generative Adversarial Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 8207–8216 (cit. on pp. 83, 93, 124, 128, 132, 141, 142).

[Sch+17] Vincent Schüssler, Eric Heitz, Johannes Hanika, and Carsten Dachsbacher. “Microfacet-based Normal Mapping for Robust Monte Carlo Path Tracing”. In: ACM Transactions on Graphics (TOG) 36.6 (2017), pp. 1–12 (cit. on p. 201).

[See66] Robert T Seeley. “Spherical Harmonics”. In: The American Mathematical Monthly 73.4P2 (1966), pp. 115–121 (cit. on pp. 185, 196).

[Sel+20] Raghavendra Selvan, Frederik Faye, Jon Middleton, and Akshay Pai. “Uncertainty Quantification in Medical Image Segmentation with Normalizing Flows”. In: Machine Learning in Medical Imaging. 2020, pp. 80–90 (cit. on p. 186).

[SKK18] Murat Sensoy, Lance Kaplan, and Melih Kandemir. “Evidential Deep Learning to Quantify Classification Uncertainty”. In: Advances in Neural Information Processing Systems 31 (2018) (cit. on pp. 126, 130).

[SDM19] Tamar Rott Shaham, Tali Dekel, and Tomer Michaeli. “SinGAN: Learning a Generative Model From a Single Natural Image”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2019 (cit. on pp. 32, 38, 57, 75, 85, 86, 90).

[Shi+20] Liang Shi, Beichen Li, Miloš Hašan, Kalyan Sunkavalli, Tamy Boubekeur, Radomir Mech, and Wojciech Matusik. “Match: Differentiable Material Graphs for Procedural Material Capture”. In: ACM Transactions on Graphics (TOG) (2020) (cit. on pp. 76, 125, 137, 138, 143, 151–153).

[Sho+19] Assaf Shocher, Shai Bagon, Phillip Isola, and Michal Irani. “InGAN: Capturing and Retargeting the "DNA" of a Natural Image”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). Oct. 2019 (cit. on p. 75).

[SCI18] Assaf Shocher, Nadav Cohen, and Michal Irani. ““Zero-Shot” Super-Resolution Using Deep Internal Learning”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2018 (cit. on p. 75).

[SK19] Connor Shorten and Taghi M Khoshgoftaar. “A Survey on Image Data Augmentation for Deep Learning”. In: Journal of Big Data 6.1 (2019), p. 60 (cit. on pp. 36, 45).

[SZ15] Karen Simonyan and Andrew Zisserman. “Very Deep Convolutional Networks for Large-Scale Image Recognition”. In: International Conference on Learning Representations (ICLR). 2015 (cit. on pp. 32, 164).

[Sin+21] Abhishek Sinha, Kumar Ayush, Jiaming Song, Burak Uzkent, Hongxia Jin, and Stefano Ermon. “Negative Data Augmentation”. In: arXiv preprint arXiv:2102.05113 (2021) (cit. on p. 93).

[Sit+20] Vincent Sitzmann, Julien Martel, Alexander Bergman, David Lindell, and Gordon Wetzstein. “Implicit Neural Representations with Periodic Activation Functions”. In: Advances in Neural Information Processing Systems 33 (2020) (cit. on pp. 75, 102, 109, 112, 120, 186, 193, 194, 262).

[Sit+21] Vincent Sitzmann, Semon Rezchikov, William T. Freeman, Joshua B. Tenenbaum, and Fredo Durand. “Light Field Networks: Neural Scene Representations with Single-Evaluation Rendering”. In: Advances in Neural Information Processing Systems. 2021 (cit. on p. 187).

[SKS02] Peter-Pike Sloan, Jan Kautz, and John Snyder. “Precomputed Radiance Transfer for Real-Time Rendering in Dynamic, Low-Frequency Lighting Environments”. In: ACM Transactions on Graphics (TOG) 21.3 (July 2002), pp. 527–536 (cit. on p. 13).

[Sne17] Xavier Snelgrove. “High-resolution Multi-scale Neural Texture Synthesis”. In: SIG-GRAPH Asia 2017 Technical Briefs. 2017, pp. 1–4 (cit. on p. 74).

[Sol+21] Ava P Soleimany, Alexander Amini, Samuel Goldman, Daniela Rus, Sangeeta N Bhatia, and Connor W Coley. “Evidential Deep Learning for Guided Molecular Property Prediction and Discovery”. In: ACS Central Science 7.8 (2021), pp. 1356–1367 (cit. on pp. 126, 135).

[Spe+22] Georg Sperl, Rosa M Sánchez-Banderas, Manwen Li, Chris Wojtan, and Miguel A Otaduy. “Estimation of Yarn-level Simulation Models for Production Fabrics”. In: ACM Transactions on Graphics (TOG) 41.4 (2022), pp. 1–15 (cit. on p. 163).

[Sri+20] Pratul P Srinivasan, Ben Mildenhall, Matthew Tancik, Jonathan T Barron, Richard Tucker, and Noah Snavely. “Lighthouse: Predicting Lighting Volumes for Spatially-Coherent Illumination”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 8080–8089 (cit. on p. 203).

[Sri+14] Nitish Srivastava, Geoffrey Hinton, Alex Krizhevsky, Ilya Sutskever, and Ruslan Salakhutdinov. “Dropout: A Simple Way to Prevent Neural Networks from Overfitting”. In: The Journal of Machine Learning Research 15.1 (2014), pp. 1929–1958 (cit. on pp. 110, 127, 141, 176).

[Ste+14] Heinz Christian Steinhausen, Dennis den Brok, Matthias B Hullin, and Reinhard Klein. “Acquiring Bidirectional Texture Functions for Large-Scale Material Samples”. In: (2014) (cit. on pp. 30, 63).

[Ste+15a] Heinz Christian Steinhausen, Dennis den Brok, Matthias B. Hullin, and Reinhard Klein. “Extrapolating Large-Scale Material BTFs under Cross-Device Constraints”. In: Vision, Modeling & Visualization. Ed. by David Bommes, Tobias Ritschel, and Thomas Schultz. The Eurographics Association, 2015, pp. 143–150 (cit. on pp. 33, 104).

[Ste+15b] Heinz Christian Steinhausen, Rodrigo Martín, Dennis den Brok, Matthias B. Hullin, and Reinhard Klein. “Extrapolation of Bidirectional Texture Functions using Texture Synthesis guided by Photometric Normals”. In: Measuring, Modeling, and Reproducing Material Appearance II (SPIE 9398). Vol. 9398. 14. San Francisco, USA, Feb. 2015 (cit. on pp. 33, 104).

[Str+22] Yannick Strümpler, Janis Postels, Ren Yang, Luc Van Gool, and Federico Tombari. “Implicit Neural Representations for Image Compression”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2022, pp. 74–91 (cit. on p. 187).

[SGK21] Vadim Sushko, Juergen Gall, and Anna Khoreva. “One-Shot GAN: Learning to Generate Samples from Single Images and Videos”. In: arXiv preprint arXiv:2103.13389 (2021) (cit. on p. 75).

[SB08] Cédric Syllebranque and Samuel Boivin. “Estimation of Mechanical Parameters of Deformable Solids from Videos”. In: The Visual Computer 24.11 (2008), pp. 963–972 (cit. on p. 158).

[Szt+21] Alejandro Sztrajman, Gilles Rainer, Tobias Ritschel, and Tim Weyrich. “Neural BRDF Representation and Importance Sampling”. In: Computer Graphics Forum. Vol. 40. 6. Wiley Online Library. 2021, pp. 332–346 (cit. on pp. 105, 111, 116, 186, 187, 193, 203, 210).

[Tak+21] Towaki Takikawa, Joey Litalien, Kangxue Yin, Karsten Kreis, Charles Loop, Derek Nowrouzezahrai, Alec Jacobson, Morgan McGuire, and Sanja Fidler. “Neural Geometric Level of Detail: Real-time Rendering with Implicit 3D Shapes”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021 (cit. on p. 186).

[Tam+11] Omer Tamuz, Ce Liu, Serge Belongie, Ohad Shamir, and Adam Tauman Kalai. “Adaptively Learning the Crowd Kernel”. In: International Conference on Machine Learning (ICML). 2011, pp. 673–680 (cit. on pp. 170, 171).

[Tan+22] Matthew Tancik, Vincent Casser, Xinchen Yan, Sabeek Pradhan, Ben Mildenhall, Pratul Srinivasan, Jonathan T. Barron, and Henrik Kretzschmar. “Block-NeRF: Scalable Large Scene Neural View Synthesis”. In: arXiv. 2022 (cit. on p. 187).

[Tan+20] Matthew Tancik, Pratul P. Srinivasan, Ben Mildenhall, Sara Fridovich-Keil, Nithin Raghavan, Utkarsh Singhal, Ravi Ramamoorthi, Jonathan T. Barron, and Ren Ng. “Fourier Features Let Networks Learn High Frequency Functions in Low Dimensional Domains”. In: Advances in Neural Information Processing Systems (2020) (cit. on pp. 75, 186, 193).

[Tew+22] Ayush Tewari, Justus Thies, Ben Mildenhall, Pratul Srinivasan, Edgar Tretschk, Wang Yifan, Christoph Lassner, Vincent Sitzmann, Ricardo Martin-Brualla, Stephen Lombardi, et al. “Advances in Neural Rendering”. In: Computer Graphics Forum. Vol. 41. 2. Wiley Online Library. 2022, pp. 703–735 (cit. on pp. 187, 212).

[Tex+20a] Ondřej Texler, David Futschik, Jakub Fišer, Michal Lukáč, Jingwan Lu, Eli Shechtman, and Daniel Sy`kora. “Arbitrary Style Transfer using Neurally-Guided Patch-Based Ssynthesis”. In: Computers & Graphics 87 (2020), pp. 62–71 (cit. on pp. 39, 45).

[Tex+20b] Ondřej Texler, David Futschik, Michal Kučera, Ondřej Jamriška, Šárka Sochorová, Menclei Chai, Sergey Tulyakov, and Daniel Sy`kora. “Interactive Video Stylization Using Few-Shot Patch-Based Training”. In: ACM Transactions on Graphics (TOG) 39.4 (2020), pp. 73–1 (cit. on pp. 30, 31, 33, 37, 39, 46, 53, 58, 61, 107, 108, 128, 129).

[TOB15] Lucas Theis, Aäron van den Oord, and Matthias Bethge. “A Note on the Evaluation of Generative Models”. In: arXiv preprint arXiv:1511.01844 (2015) (cit. on p. 41).

[Tom94] Shoji Tominaga. “Dichromatic reflection models for a variety of materials”. In: Color Research & Application 19.4 (1994), pp. 277–285 (cit. on p. 10).

[Ton+02] Xin Tong, Jingdan Zhang, Ligang Liu, Xi Wang, Baining Guo, and Heung-Yeung Shum. “Synthesis of Bidirectional Texture Functions on Arbitrary Surfaces”. In: ACM Transactions on Graphics (TOG) 21.3 (2002), pp. 665–672 (cit. on p. 103).

[TS67] Kenneth E Torrance and Ephraim M Sparrow. “Theory for Off-Specular Reflection from Roughened Surfaces”. In: Josa 57.9 (1967), pp. 1105–1114 (cit. on p. 11).

[Tu+20] Peihan Tu, Li-Yi Wei, Koji Yatani, Takeo Igarashi, and Matthias Zwicker. “Continuous Curve Textures”. In: ACM Transactions on Graphics (TOG) 39.6 (2020), pp. 1–16 (cit. on p. 72).

[TK74] Amos Tversky and Daniel Kahneman. “Judgment under Uncertainty: Heuristics and Biases: Biases in Judgments Reveal Some Heuristics of Thinking Under Uncertainty.” In: Science 185.4157 (1974), pp. 1124–1131 (cit. on p. 169).

[Uly+16] Dmitry Ulyanov, Vadim Lebedev, Andrea Vedaldi, and Victor S Lempitsky. “Texture Networks: Feed-Forward Synthesis of Textures and Stylized Images.” In: International Conference of Machine Learning (ICML). Vol. 1. 2. 2016, p. 4 (cit. on p. 76).

[UVL18] Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. “Deep Image Prior”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2018 (cit. on p. 75).

[UVL16] Dmitry Ulyanov, Andrea Vedaldi, and Victor Lempitsky. “Instance Normalization: The Missing Ingredient for Fast Stylization”. In: arXiv preprint arXiv:1607.08022 (2016) (cit. on pp. 77, 94).

[Van11] Dietger G Van Antwerpen. “Unbiased Physically Nased Rendering on the GPU”. MA thesis. Electrical Engineering, Mathematics and Computer Science, 2011 (cit. on p. 195).

[Vas+17] Ashish Vaswani, Noam Shazeer, Niki Parmar, Jakob Uszkoreit, Llion Jones, Aidan N Gomez, Łukasz Kaiser, and Illia Polosukhin. “Attention Is All You Need”. In: Advances in Neural Information Processing Systems 30 (2017) (cit. on p. 164).

[VG95] Eric Veach and Leonidas J. Guibas. “Optimally Combining Sampling Techniques for Monte Carlo Rendering”. In: Proceedings of the 22nd Annual Conference on Computer Graphics and Interactive Techniques. SIGGRAPH ’95. Association for Computing Machinery, 1995 (cit. on p. 184).

[VPS21] Giuseppe Vecchio, Simone Palazzo, and Concetto Spampinato. “SurfaceNet: Adversarial SVBRDF Estimation from a Single Image”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 12840–12848 (cit. on pp. 17, 76, 123, 125, 128, 129, 132, 258).

[Ver+22] Dor Verbin, Peter Hedman, Ben Mildenhall, Todd Zickler, Jonathan T. Barron, and Pratul P. Srinivasan. “Ref-NeRF: Structured View-Dependent Appearance for Neural Radiance Fields”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2022 (cit. on pp. 187, 212).

[VZH20] Yael Vinker, Nir Zabari, and Yedid Hoshen. “Training End-to-end Single Image Generators without GANs”. In: arXiv preprint arXiv:2004.06014 (2020) (cit. on p. 75).

[VMF09] Pascal Volino, Nadia Magnenat-Thalmann, and Francois Faure. “A Simple Approach to Nonlinear Tensile Stiffness for Accurate Cloth Simulation”. In: ACM Transactions on Graphics (TOG) 28.4 (2009), Article–No (cit. on pp. 158, 161, 162).

[Wal+14] Ingo Wald, Sven Woop, Carsten Benthin, Gregory S. Johnson, and Manfred Ernst. “Embree: A Kernel Framework for Efficient CPU Ray Tracing”. In: ACM Transactions on Graphics (TOG) 33.4 (2014) (cit. on p. 195).

[Wal+07] Bruce Walter, Stephen R Marschner, Hongsong Li, and Kenneth E Torrance. “Microfacet Models for Refraction through Rough Surfaces”. In: Proceedings of the 18th Eurographics Conference on Rendering Techniques. 2007, pp. 195–206 (cit. on p. 127).

[WZY22] Cairong Wang, Yiming Zhu, and Chun Yuan. “Diverse Image Inpainting with Normalizing Flow”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2022, pp. 53–69 (cit. on p. 186).

[Wan+22a] Can Wang, Menglei Chai, Mingming He, Dongdong Chen, and Jing Liao. “Clip-nerf: Text-and-image Driven Manipulation of Neural Radiance Fields”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 3835–3844 (cit. on p. 116).

[Wan+22b] Chen Wang, Xiang Wang, Jiawei Zhang, Liang Zhang, Xiao Bai, Xin Ning, Jun Zhou, and Edwin Hancock. “Uncertainty Estimation for Stereo Matching Based on Evidential Deep Learning”. In: Pattern Recognition 124 (2022), p. 108498 (cit. on p. 130).

[Wan+18] Guotai Wang, Wenqi Li, Maria A Zuluaga, Rosalind Pratt, Premal A Patel, Michael Aertsen, Tom Doel, Anna L David, Jan Deprest, Sébastien Ourselin, et al. “Interactive Medical Image Segmentation using Deep Learning with Image-specific Fine Tuning”. In: IEEE transactions on medical imaging 37.7 (2018), pp. 1562–1573 (cit. on p. 159).

[WOR11] Huamin Wang, James F O’Brien, and Ravi Ramamoorthi. “Data-driven Elastic Models for Cloth: Modeling and Measurement”. In: ACM Transactions on Graphics (TOG) 30.4 (2011), pp. 1–12 (cit. on p. 158).

[Wan+09] Jiaping Wang, Peiran Ren, Minmin Gong, John Snyder, and Baining Guo. “All-Frequency Rendering of Dynamic, Spatially-Varying Reflectance”. In: 28.5 (Dec. 2009), pp. 1–10 (cit. on p. 13).

[Wan+22c] Jiayi Wang, Diogo Luvizon, Franziska Mueller, Florian Bernard, Adam Kortylewski, Dan Casas, and Christian Theobalt. “HandFlow: Quantifying View-Dependent 3D Ambiguity in Two-Hand Reconstruction with Normalizing Flow”. In: International Symposium on Vision, Modeling, and Visualization. 2022 (cit. on p. 186).

[Wan+20a] Jiayun Wang, Yubei Chen, Rudrasis Chakraborty, and Stella X. Yu. “Orthogonal Convolutional Neural Networks”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). June 2020 (cit. on p. 93).

[Wan+22d] Liao Wang, Jiakai Zhang, Xinhang Liu, Fuqiang Zhao, Yanshun Zhang, Yingliang Zhang, Minye Wu, Jingyi Yu, and Lan Xu. “Fourier Plenoctrees for Dynamic Radiance Field Rendering in Real-time”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2022, pp. 13524–13534 (cit. on pp. 203, 212).

[Wan+20b] Sinong Wang, Belinda Z Li, Madian Khabsa, Han Fang, and Hao Ma. “Linformer: Self-attention with Linear Complexity”. In: arXiv preprint arXiv:2006.04768 (2020) (cit. on pp. 124, 127, 141, 144).

[Wan+19] Ting-Chun Wang, Ming-Yu Liu, Andrew Tao, Guilin Liu, Bryan Catanzaro, and Jan Kautz. “Few-shot Video-to-Video Synthesis”. In: Advances in Neural Information Processing Systems. 2019, pp. 5013–5024 (cit. on p. 37).

[Wan+20c] Xiaofei Wang, Yiwen Han, Victor CM Leung, Dusit Niyato, Xueqiang Yan, and Xu Chen. “Convergence of Edge Computing and Deep Learning: A Comprehensive Survey”. In: IEEE Communications Surveys & Tutorials 22.2 (2020), pp. 869–904 (cit. on p. 212).

[Wan+20d] Yaqing Wang, Quanming Yao, James T Kwok, and Lionel M Ni. “Generalizing From a Few Examples: A Survey on Few-Shot Learning”. In: ACM Computing Surveys (CSUR) 53.3 (2020), pp. 1–34 (cit. on p. 37).

[Wan+04] Zhou Wang, Alan C Bovik, Hamid R Sheikh, and Eero P Simoncelli. “Image Quality Assessment: From Error Visibility to Structural Similarity”. In: IEEE Transactions on Image Processing 13.4 (2004), pp. 600–612 (cit. on pp. 50, 54, 85, 86, 90, 92, 160, 198, 199, 202).

[Wan+21] Zian Wang, Jonah Philion, Sanja Fidler, and Jan Kautz. “Learning Indoor Inverse Rendering with 3D Spatially-Varying Lighting”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV). 2021, pp. 12538–12547 (cit. on p. 203).

[Wan+23] Zian Wang, Tianchang Shen, Jun Gao, Shengyu Huang, Jacob Munkberg, Jon Hasselgren, Zan Gojcic, Wenzheng Chen, and Sanja Fidler. “Neural Fields meet Explicit Geometric Representations for Inverse Rendering of Urban Scenes”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2023 (cit. on pp. 203, 212).

[WGK14] Michael Weinmann, Juergen Gall, and Reinhard Klein. “Material Classification Based on Training Data Synthesized Using a BTF Database”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer International Publishing, 2014, pp. 156–171 (cit. on pp. 50, 103, 108, 111, 151).

[WKW16] Karl Weiss, Taghi M Khoshgoftaar, and DingDing Wang. “A Survey of Transfer Learning”. In: Journal of Big data 3.1 (2016), pp. 1–40 (cit. on p. 159).

[Wen+22] Tao Wen, Beibei Wang, Lei Zhang, Jie Guo, and Nicolas Holzschuch. “SVBRDF Recovery from a Single Image with Highlights Using a Pre-trained Generative Adversarial Network”. In: Computer Graphics Forum. Wiley Online Library. 2022 (cit. on p. 125).

[Woo+18] Sanghyun Woo, Jongchan Park, Joon-Young Lee, and In So Kweon. “Cbam: Convolutional block attention module”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 3–19 (cit. on pp. 109, 119, 128, 141, 145, 164).

[WTX22] Qiling Wu, Jianchao Tan, and Kun Xu. “PaletteNeRF: Palette-based Color Editing for NeRFs”. In: arXiv preprint arXiv:2212.12871 (2022) (cit. on p. 116).

[WH18] Yuxin Wu and Kaiming He. “Group Normalization”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2018, pp. 3–19 (cit. on pp. 64, 67, 69, 127, 141, 144).

[Xie+22] Yiheng Xie, Towaki Takikawa, Shunsuke Saito, Or Litany, Shiqin Yan, Numair Khan, Federico Tombari, James Tompkin, Vincent Sitzmann, and Srinath Sridhar. “Neural Fields in Visual Computing and Beyond”. In: Computer Graphics Forum (2022) (cit. on pp. 187, 212).

[Xu+13] Kun Xu, Wei-Lun Sun, Zhao Dong, Dan-Yong Zhao, Run-Dong Wu, and Shi-Min Hu. “Anisotropic Spherical Gaussians”. In: ACM Transactions on Graphics (TOG) 32.6 (2013), pp. 1–11 (cit. on pp. 185, 196, 199, 202).

[Yan+21] Guandao Yang, Serge Belongie, Bharath Hariharan, and Vladlen Koltun. “Geometry processing with Neural Fields”. In: Advances in Neural Information Processing Systems. Vol. 34. 2021, pp. 22483–22497 (cit. on p. 186).

[YLL17] Shan Yang, Junbang Liang, and Ming C Lin. “Learning-based Cloth Material Recovery from Video”. In: Proceedings of the IEEE International Conference on Computer Vision (ICCV). 2017, pp. 4383–4393 (cit. on pp. 156, 159).

[YL15] Shan Yang and Ming C Lin. “Materialcloning: Acquiring Elasticity Parameters from Images for Medical Applications”. In: IEEE Transactions on Visualization and Computer Graphics 22.9 (2015), pp. 2122–2135 (cit. on p. 158).

[Yan+18] Shan Yang, Zherong Pan, Tanya Amert, Ke Wang, Licheng Yu, Tamara Berg, and Ming C Lin. “Physics-inspired Garment Recovery from a Single-view Image”. In: ACM Transactions on Graphics (TOG) 37.5 (2018), pp. 1–14 (cit. on p. 158).

[Yao+22] Yao Yao, Jingyang Zhang, Jingbo Liu, Yihang Qu, Tian Fang, David McKinnon, Yanghai Tsin, and Long Quan. “NeILF: Neural Incident Light Field for Material and Lighting Estimation”. In: Proceedings of the European Conference on Computer Vision (ECCV). 2022 (cit. on p. 187).

[Ye+22] Weicai Ye, Shuo Chen, Chong Bao, Hujun Bao, Marc Pollefeys, Zhaopeng Cui, and Guofeng Zhang. “Intrinsicnerf: Learning Intrinsic Neural Radiance Fields for Editable Novel View Synthesis”. In: arXiv preprint arXiv:2210.00647 (2022) (cit. on p. 116).

[Ye+21] Wenjie Ye, Yue Dong, Pieter Peers, and Baining Guo. “Deep Reflectance Scanning: Recovering Spatially-varying Material Appearance from a Flash-lit Video Sequence”. In: Computer Graphics Forum. Vol. 40. 6. Wiley Online Library. 2021, pp. 409–427 (cit. on p. 125).

[Ye+18] Wenjie Ye, Xiao Li, Yue Dong, Pieter Peers, and Xin Tong. “Single Image Surface Appearance Modeling with Self-Augmented CNNs and Inexact Supervision”. In: Computer Graphics Forum. Vol. 37. 7. Wiley Online Library. 2018, pp. 201–211 (cit. on pp. 34, 125).

[Yep20] Tyler Yep. torchinfo. Mar. 2020 (cit. on p. 113).

[YS19] Ye Yu and William AP Smith. “InverseRenderNet: Learning Single Image Inverse Rendering”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 3155–3164 (cit. on pp. 78, 203).

[YS21] Ye Yu and William AP Smith. “Outdoor Inverse Rendering from a Single Image using Multiview Self-supervision”. In: IEEE Transactions on Pattern Analysis and Machine Intelligence 44.7 (2021), pp. 3659–3675 (cit. on p. 203).

[Yun+19] Sangdoo Yun, Dongyoon Han, Seong Joon Oh, Sanghyuk Chun, Junsuk Choe, and Youngjoon Yoo. “Cutmix: Regularization strategy to train strong classifiers with localizable features”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 6023–6032 (cit. on p. 93).

[Zha+19a] Bo Zhang, Mingming He, Jing Liao, Pedro V Sander, Lu Yuan, Amine Bermak, and Dong Chen. “Deep Exemplar-based Video Colorization”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 8052–8061 (cit. on p. 33).

[Zha+19b] Han Zhang, Ian Goodfellow, Dimitris Metaxas, and Augustus Odena. “Self-attention Generative Adversarial Networks”. In: International Conference on Machine Learning (ICML). PMLR. 2019, pp. 7354–7363 (cit. on pp. 164, 166, 176).

[ZDN16] Hang Zhang, Kristin Dana, and Ko Nishino. “Friction from Reflectance: Deep Reflectance Codes for Predicting Physical Surface Properties from One-shot In-field Reflectance”. In: Proceedings of the European Conference on Computer Vision (ECCV). Springer. 2016, pp. 808–824 (cit. on p. 159).

[Zha+20] Jingzhao Zhang, Tianxing He, Suvrit Sra, and Ali Jadbabaie. “Why Gradient Clipping Accelerates Training: A Theoretical Justification for Adaptivity”. In: International Conference on Learning Representations (ICLR). 2020 (cit. on pp. 110, 194).

[Zha+21] Kai Zhang, Fujun Luan, Qianqian Wang, Kavita Bala, and Noah Snavely. “Physg: Inverse Rendering with Spherical Gaussians for Physics-based Material Editing and Relighting”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2021, pp. 5453–5462 (cit. on p. 203).

[ZL12] Lin Zhang and Hongyu Li. “SR-SIM: A Fast and High Performance IQA Index Based on Spectral Residual”. In: IEEE International Conference on Image Processing. IEEE. 2012, pp. 1473–1476 (cit. on p. 160).

[ZSL14] Lin Zhang, Ying Shen, and Hongyu Li. “VSI: A Visual Saliency-induced Index for Perceptual Image Quality Assessment”. In: IEEE Transactions on Image processing 23.10 (2014), pp. 4270–4281 (cit. on p. 160).

[Zha+11] Lin Zhang, Lei Zhang, Xuanqin Mou, and David Zhang. “FSIM: A Feature Similarity Index for Image Quality Assessment”. In: IEEE Transactions on Image Processing 20.8 (2011), pp. 2378–2386 (cit. on p. 160).

[Zha+18] Richard Zhang, Phillip Isola, Alexei A. Efros, Eli Shechtman, and Oliver Wang. “The Unreasonable Effectiveness of Deep Features as a Perceptual Metric”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2018, pp. 586–595 (cit. on pp. 32, 39, 41, 49, 50, 54, 57, 67, 85, 86, 90, 110, 128, 142, 157, 160, 169, 170, 178, 198, 199, 202).

[ZJK20] Hengshuang Zhao, Jiaya Jia, and Vladlen Koltun. “Exploring Self-attention for Image Recognition”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2020, pp. 10076–10085 (cit. on p. 164).

[Zho+20] Zhun Zhong, Liang Zheng, Guoliang Kang, Shaozi Li, and Yi Yang. “Random Erasing Data Augmentation”. In: Proceedings of the AAAI Conference on Artificial Intelligence (AAAI). 2020 (cit. on pp. 111, 129, 142, 176).

[Zho+16] Bolei Zhou, Aditya Khosla, Agata Lapedriza, Aude Oliva, and Antonio Torralba. “Learning Deep Features for Discriminative Localization”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2016, pp. 2921–2929 (cit. on p. 164).

[Zho+22] Xilong Zhou, Miloš Hašan, Valentin Deschaintre, Paul Guerrero, Kalyan Sunkavalli, and Nima Kalantari. “TileGen: Tileable, Controllable Material Generation and Capture”. In: ACM Transactions on Graphics (Proc. SIGGRAPH Asia) (2022) (cit. on pp. 105, 107, 109, 116, 123).

[Zho+23] Xilong Zhou, Miloš Hašan, Valentin Deschaintre, Paul Guerrero, Kalyan Sunkavalli, and Nima Khademi Kalantari. “A Semi-Procedural Convolutional Material Prior”. In: Computer Graphics Forum. Wiley Online Library. 2023 (cit. on p. 105).

[ZK21] Xilong Zhou and Nima Khademi Kalantari. “Adversarial Single-Image SVBRDF Estimation with Hybrid Training”. In: Computer Graphics Forum. Vol. 40. 2. Wiley Online Library. 2021, pp. 315–325 (cit. on pp. 123, 125, 127, 137, 138, 141, 152, 153).

[Zho+17] Yang Zhou, Huajie Shi, Dani Lischinski, Minglun Gong, Johannes Kopf, and Hui Huang. “Analysis and Controlled Synthesis of Inhomogeneous Textures”. In: Computer Graphics Forum. Vol. 36. 2. 2017, pp. 199–212 (cit. on p. 74).

[Zho+18] Yang Zhou, Zhen Zhu, Xiang Bai, Dani Lischinski, Daniel Cohen-Or, and Hui Huang. “Non-Stationary Texture Synthesis by Adversarial Expansion”. In: ACM Transactions on Graphics (TOG) 37.4 (July 2018) (cit. on pp. 34, 38, 56, 72, 74, 76, 77, 79, 84, 85, 87, 90, 94, 104).

[Zho+19] Yizhou Zhou, Xiaoyan Sun, Zheng-Jun Zha, and Wenjun Zeng. “Context-Reinforced Semantic Segmentation”. In: Proceedings of the IEEE/CVF Conference on Computer Vision and Pattern Recognition (CVPR). 2019, pp. 4046–4055 (cit. on p. 41).

[Zhu+17a] Jun Yan Zhu, Taesung Park, Phillip Isola, and Alexei A. Efros. “Unpaired Image-to-Image Translation Using Cycle-Consistent Adversarial Networks”. In: Proceedings of the IEEE/CVF International Conference on Computer Vision (ICCV) 2017-Octob (Mar. 2017), pp. 2242–2251 (cit. on pp. 78, 79, 83, 94).

[Zhu+17b] Jun-Yan Zhu, Richard Zhang, Deepak Pathak, Trevor Darrell, Alexei A Efros, Oliver Wang, and Eli Shechtman. “Toward Multimodal Image-to-Image Translation”. In: Advances in Neural Information Processing Systems. 2017, pp. 465–476 (cit. on p. 37).

[Zhu+21] Wei Zhu, Xian Guo, Dai Owaki, Kyo Kutsuzawa, and Mitsuhiro Hayashibe. “A Survey of Sim-to-real Transfer Techniques Applied to Reinforcement Learning for Bioinspired Robots”. In: IEEE Transactions on Neural Networks and Learning Systems (2021) (cit. on pp. 156, 212).

_______________________________

1 Additional publication details and results are included on the project website.