allow for training the Prior network with precomputed CLIP embeddings (or text encodings)

refactor so that the causal transformer in the diffusion prior network can be conditioned without text encodings (for Laions parallel efforts, although it seems from the paper it is needed)
make sure non-latent diffusion still works
2026-02-12 11:34:29 +01:00 · 2022-04-26 09:29:51 -07:00 · 2022-04-26 09:00:11 -07:00 · 2022-04-26 08:36:00 -07:00 · 2022-04-26 07:39:04 -07:00 · 2022-04-25 19:24:13 -07:00
6 changed files with 343 additions and 55 deletions
--- a/README.md
+++ b/README.md
@@ -12,7 +12,7 @@ This model is SOTA for text-to-image for now.

 Please join <a href="https://discord.gg/xBPBXfcFHd"><img alt="Join us on Discord" src="https://img.shields.io/discord/823813159592001537?color=5865F2&logo=discord&logoColor=white"></a> if you are interested in helping out with the replication

-There was enough interest for a Jax version. It will be completed after the Pytorch version shows signs of life on my toy tasks. <a href="https://github.com/lucidrains/dalle2-jax">Placeholder repository</a>. I will also eventually extend this to <a href="https://github.com/lucidrains/dalle2-video">text to video</a>, once the repository is in a good place.
+There was enough interest for a <a href="https://github.com/lucidrains/dalle2-jax">Jax version</a>. I will also eventually extend this to <a href="https://github.com/lucidrains/dalle2-video">text to video</a>, once the repository is in a good place.

 ## Install

@@ -246,13 +246,6 @@ loss = decoder(images, unet_number = 2)
 loss.backward()

 # do the above for many steps for both unets
-
-# then it will learn to generate images based on the CLIP image embeddings
-
-# chaining the unets from lowest resolution to highest resolution (thus cascading)
-
-mock_image_embed = torch.randn(1, 512).cuda()
-images = decoder.sample(mock_image_embed) # (1, 3, 512, 512)
 ```

 Finally, to generate the DALL-E2 images from text. Insert the trained `DiffusionPrior` as well as the `Decoder` (which wraps `CLIP`, the causal transformer, and unet(s))
@@ -383,6 +376,75 @@ You can also train the decoder on images of greater than the size (say 512x512)

 For the layperson, no worries, training will all be automated into a CLI tool, at least for small scale training.

+## Training on Preprocessed CLIP Embeddings
+
+It is likely, when scaling up, that you would first preprocess your images and text into corresponding embeddings before training the prior network. You can do so easily by simply passing in `image_embed`, `text_embed`, and optionally `text_encodings` and `text_mask`
+
+Working example below
+
+```python
+import torch
+from dalle2_pytorch import DiffusionPriorNetwork, DiffusionPrior, CLIP
+
+# get trained CLIP from step one
+
+clip = CLIP(
+    dim_text = 512,
+    dim_image = 512,
+    dim_latent = 512,
+    num_text_tokens = 49408,
+    text_enc_depth = 6,
+    text_seq_len = 256,
+    text_heads = 8,
+    visual_enc_depth = 6,
+    visual_image_size = 256,
+    visual_patch_size = 32,
+    visual_heads = 8,
+).cuda()
+
+# setup prior network, which contains an autoregressive transformer
+
+prior_network = DiffusionPriorNetwork(
+    dim = 512,
+    depth = 6,
+    dim_head = 64,
+    heads = 8
+).cuda()
+
+# diffusion prior network, which contains the CLIP and network (with transformer) above
+
+diffusion_prior = DiffusionPrior(
+    net = prior_network,
+    clip = clip,
+    timesteps = 100,
+    cond_drop_prob = 0.2,
+    condition_on_text_encodings = False  # this probably should be true, but just to get Laion started
+).cuda()
+
+# mock data
+
+text = torch.randint(0, 49408, (4, 256)).cuda()
+images = torch.randn(4, 3, 256, 256).cuda()
+
+# precompute the text and image embeddings
+# here using the diffusion prior class, but could be done with CLIP alone
+
+clip_image_embeds = diffusion_prior.get_image_embed(images)
+clip_text_embeds = diffusion_prior.get_text_cond(text).get('text_embed')
+
+# feed text and images into diffusion prior network
+
+loss = diffusion_prior(
+    text_embed = clip_text_embeds,
+    image_embed = clip_image_embeds
+)
+
+loss.backward()
+
+# do the above for many many many steps
+# now the diffusion prior can generate image embeddings from the text embeddings
+```
+
 ## Experimental

 ### DALL-E2 with Latent Diffusion
@@ -530,9 +592,11 @@ Once built, images will be saved to the same directory the command is invoked
 - [x] offload unets not being trained on to CPU for memory efficiency (for training each resolution unets separately)
 - [x] build out latent diffusion architecture, with the vq-reg variant (vqgan-vae), make it completely optional and compatible with cascading ddpms
 - [x] for decoder, allow ability to customize objective (predict epsilon vs x0), in case latent diffusion does better with prediction of x0
+- [x] use attention-based upsampling https://arxiv.org/abs/2112.11435
 - [ ] spend one day cleaning up tech debt in decoder
 - [ ] become an expert with unets, cleanup unet code, make it fully configurable, port all learnings over to https://github.com/lucidrains/x-unet
 - [ ] copy the cascading ddpm code to a separate repo (perhaps https://github.com/lucidrains/denoising-diffusion-pytorch) as the main contribution of dalle2 really is just the prior network
+- [ ] transcribe code to Jax, which lowers the activation energy for distributed training, given access to TPUs
 - [ ] train on a toy task, offer in colab
 - [ ] extend diffusion head to use diffusion-gan (potentially using lightweight-gan) to speed up inference
 - [ ] bring in tools to train vqgan-vae
@@ -568,7 +632,7 @@ Once built, images will be saved to the same directory the command is invoked

 ```bibtex
@inproceedings{Liu2022ACF,
-    title   = {A ConvNet for the 2020s},
+    title   = {A ConvNet for the 2020https://arxiv.org/abs/2112.11435s},
    author  = {Zhuang Liu and Hanzi Mao and Chaozheng Wu and Christoph Feichtenhofer and Trevor Darrell and Saining Xie},
    year    = {2022}
 }
--- a/dalle2_pytorch/attention.py
+++ b/dalle2_pytorch/attention.py
@@ -0,0 +1,125 @@
+import torch
+from torch import nn, einsum
+import torch.nn.functional as F
+
+from einops import rearrange, repeat
+
+class LayerNormChan(nn.Module):
+    def __init__(
+        self,
+        dim,
+        eps = 1e-5
+    ):
+        super().__init__()
+        self.eps = eps
+        self.gamma = nn.Parameter(torch.ones(1, dim, 1, 1))
+
+    def forward(self, x):
+        var = torch.var(x, dim = 1, unbiased = False, keepdim = True)
+        mean = torch.mean(x, dim = 1, keepdim = True)
+        return (x - mean) / (var + self.eps).sqrt() * self.gamma
+
+# attention-based upsampling
+# from https://arxiv.org/abs/2112.11435
+
+class QueryAndAttend(nn.Module):
+    def __init__(
+        self,
+        *,
+        dim,
+        num_queries = 1,
+        dim_head = 32,
+        heads = 8,
+        window_size = 3
+    ):
+        super().__init__()
+        self.scale = dim_head ** -0.5
+        inner_dim = dim_head * heads
+        self.heads = heads
+        self.dim_head = dim_head
+        self.window_size = window_size
+        self.num_queries = num_queries
+
+        self.rel_pos_bias = nn.Parameter(torch.randn(heads, num_queries, window_size * window_size, 1, 1))
+
+        self.queries = nn.Parameter(torch.randn(heads, num_queries, dim_head))
+        self.to_kv = nn.Conv2d(dim, dim_head * 2, 1, bias = False)
+        self.to_out = nn.Conv2d(inner_dim, dim, 1, bias = False)
+
+    def forward(self, x):
+        """
+        einstein notation
+        b - batch
+        h - heads
+        l - num queries
+        d - head dimension
+        x - height
+        y - width
+        j - source sequence for attending to (kernel size squared in this case)
+        """
+
+        wsz, heads, dim_head, num_queries = self.window_size, self.heads, self.dim_head, self.num_queries
+        batch, _, height, width = x.shape
+
+        is_one_query = self.num_queries == 1
+
+        # queries, keys, values
+
+        q = self.queries * self.scale
+        k, v = self.to_kv(x).chunk(2, dim = 1)
+
+        # similarities
+
+        sim = einsum('h l d, b d x y -> b h l x y', q, k)
+        sim = rearrange(sim, 'b ... x y -> b (...) x y')
+
+        # unfold the similarity scores, with float(-inf) as padding value
+
+        mask_value = -torch.finfo(sim.dtype).max
+        sim = F.pad(sim, ((wsz // 2,) * 4), value = mask_value)
+        sim = F.unfold(sim, kernel_size = wsz)
+        sim = rearrange(sim, 'b (h l j) (x y) -> b h l j x y', h = heads, l = num_queries, x = height, y = width)
+
+        # rel pos bias
+
+        sim = sim + self.rel_pos_bias
+
+        # numerically stable attention
+
+        sim = sim - sim.amax(dim = -3, keepdim = True).detach()
+        attn = sim.softmax(dim = -3)
+
+        # unfold values
+
+        v = F.pad(v, ((wsz // 2,) * 4), value = 0.)
+        v = F.unfold(v, kernel_size = wsz)
+        v = rearrange(v, 'b (d j) (x y) -> b d j x y', d = dim_head, x = height, y = width)
+
+        # aggregate values
+
+        out = einsum('b h l j x y, b d j x y -> b l h d x y', attn, v)
+
+        # combine heads
+
+        out = rearrange(out, 'b l h d x y -> (b l) (h d) x y')
+        out = self.to_out(out)
+        out = rearrange(out, '(b l) d x y -> b l d x y', b = batch)
+
+        # return original input if one query
+
+        if is_one_query:
+            out = rearrange(out, 'b 1 ... -> b ...')
+
+        return out
+
+class QueryAttnUpsample(nn.Module):
+    def __init__(self, dim, **kwargs):
+        super().__init__()
+        self.norm = LayerNormChan(dim)
+        self.qna = QueryAndAttend(dim = dim, num_queries = 4, **kwargs)
+
+    def forward(self, x):
+        x = self.norm(x)
+        out = self.qna(x)
+        out = rearrange(out, 'b (w1 w2) c h w -> b c (h w1) (w w2)', w1 = 2, w2 = 2)
+        return out
--- a/dalle2_pytorch/dalle2_pytorch.py
+++ b/dalle2_pytorch/dalle2_pytorch.py
@@ -17,6 +17,7 @@ from kornia.filters import gaussian_blur2d

 from dalle2_pytorch.tokenizer import tokenizer
 from dalle2_pytorch.vqgan_vae import NullVQGanVAE, VQGanVAE
+from dalle2_pytorch.attention import QueryAttnUpsample

 # use x-clip

@@ -420,25 +421,41 @@ class DiffusionPriorNetwork(nn.Module):
        image_embed,
        diffusion_timesteps,
        *,
-        text_encodings,
        text_embed,
+        text_encodings = None,
        mask = None,
        cond_drop_prob = 0.2
    ):
-        batch, text_enc_len, device = image_embed.shape[0], text_encodings.shape[-2], image_embed.device
+        batch, dim, device, dtype = *image_embed.shape, image_embed.device, image_embed.dtype

        # in section 2.2, last paragraph
        # "... consisting of encoded text, CLIP text embedding, diffusion timestep embedding, noised CLIP image embedding, final embedding for prediction"

        text_embed, image_embed = rearrange_many((text_embed, image_embed), 'b d -> b 1 d')

+        # make text encodings optional
+        # although the paper seems to suggest it is present <--
+
+        if not exists(text_encodings):
+            text_encodings = torch.empty((batch, 0, dim), device = device, dtype = dtype)
+
+        if not exists(mask):
+            mask = torch.ones((batch, text_encodings.shape[-2]), device = device, dtype = torch.bool)
+
+        # classifier free guidance
+
+        cond_prob_mask = prob_mask_like((batch,), cond_drop_prob, device = device)
+        cond_prob_mask = rearrange(cond_prob_mask, 'b -> b 1')
+
+        mask &= cond_prob_mask
+
+        # whether text embedding is masked or not depends on the classifier free guidance conditional masking
+
+        mask = torch.cat((mask, cond_prob_mask), dim = 1)
+
        # whether text embedding is used for conditioning depends on whether text encodings are available for attention (for classifier free guidance, even though it seems from the paper it was not used in the prior ddpm, as the objective is different)
        # but let's just do it right

-        if exists(mask):
-            not_all_masked_out = mask.any(dim = -1)
-            mask = torch.cat((mask, rearrange(not_all_masked_out, 'b -> b 1')), dim = 1)
-
        if exists(mask):
            mask = F.pad(mask, (0, 2), value = True) # extend mask for text embedding, noised image embedding, time step embedding, and learned query

@@ -454,16 +471,6 @@ class DiffusionPriorNetwork(nn.Module):
            learned_queries
        ), dim = -2)

-        # mask if it doesn't exist
-
-        if not exists(mask):
-            mask = torch.ones((batch, text_enc_len), device = device, dtype = torch.bool)
-
-        # classifier free guidance
-
-        cond_prob_mask = prob_mask_like((batch,), cond_drop_prob, device = device)
-        mask &= rearrange(cond_prob_mask, 'b -> b 1')
-
        # attend

        tokens = self.causal_transformer(tokens, mask = mask)
@@ -483,8 +490,9 @@ class DiffusionPrior(nn.Module):
        timesteps = 1000,
        cond_drop_prob = 0.2,
        loss_type = "l1",
-        predict_x0 = True,
+        predict_x_start = True,
        beta_schedule = "cosine",
+        condition_on_text_encodings = True, # the paper suggests this is needed, but you can turn it off for your CLIP preprocessed text embed -> image embed training
    ):
        super().__init__()
        assert isinstance(clip, CLIP)
@@ -495,9 +503,11 @@ class DiffusionPrior(nn.Module):
        self.image_embed_dim = clip.dim_latent
        self.channels = clip.image_channels
        self.image_size = clip.image_size
-        self.cond_drop_prob = cond_drop_prob

-        self.predict_x0 = predict_x0
+        self.cond_drop_prob = cond_drop_prob
+        self.condition_on_text_encodings = condition_on_text_encodings
+
+        self.predict_x_start = predict_x_start
        # in paper, they do not predict the noise, but predict x0 directly for image embedding, claiming empirically better results. I'll just offer both.

        if beta_schedule == "cosine":
@@ -560,6 +570,10 @@ class DiffusionPrior(nn.Module):
        text_cls, text_encodings = text_encodings[:, 0], text_encodings[:, 1:]
        text_embed = self.clip.to_text_latent(text_cls)
        text_embed = l2norm(text_embed)
+
+        if not self.condition_on_text_encodings:            
+            return dict(text_embed = text_embed)
+
        return dict(text_encodings = text_encodings, text_embed = text_embed, mask = text != 0)

    def q_mean_variance(self, x_start, t):
@@ -586,14 +600,14 @@ class DiffusionPrior(nn.Module):
    def p_mean_variance(self, x, t, text_cond, clip_denoised: bool):
        pred = self.net(x, t, **text_cond)

-        if self.predict_x0:
+        if self.predict_x_start:
            x_recon = pred
            # not 100% sure of this above line - for any spectators, let me know in the github issues (or through a pull request) if you know how to correctly do this
            # i'll be rereading https://arxiv.org/abs/2111.14822, where i think a similar approach is taken
        else:
            x_recon = self.predict_start_from_noise(x, t = t, noise = pred)

-        if clip_denoised and not self.predict_x0:
+        if clip_denoised and not self.predict_x_start:
            x_recon.clamp_(-1., 1.)

        model_mean, posterior_variance, posterior_log_variance = self.q_posterior(x_start=x_recon, x_t=x, t=t)
@@ -639,7 +653,7 @@ class DiffusionPrior(nn.Module):
            **text_cond
        )

-        to_predict = noise if not self.predict_x0 else image_embed
+        to_predict = noise if not self.predict_x_start else image_embed

        if self.loss_type == 'l1':
            loss = F.l1_loss(to_predict, x_recon)
@@ -678,13 +692,41 @@ class DiffusionPrior(nn.Module):
        top_image_embeds = image_embeds.gather(1, top_sim_indices)
        return rearrange(top_image_embeds, 'b 1 d -> b d')

-    def forward(self, text, image, *args, **kwargs):
-        b, device, img_size, = image.shape[0], image.device, self.image_size
-        check_shape(image, 'b c h w', h = img_size, w = img_size, c = self.channels)
+    def forward(
+        self,
+        text = None,
+        image = None,
+        text_embed = None,      # allow for training on preprocessed CLIP text and image embeddings
+        image_embed = None,
+        text_encodings = None,  # as well as CLIP text encodings
+        text_mask = None,       # text mask <- may eventually opt for the learned padding tokens technique from DALL-E1 to reduce complexity
+        *args,
+        **kwargs
+    ):
+        assert exists(text) ^ exists(text_embed), 'either text or text embedding must be supplied'
+        assert exists(image) ^ exists(image_embed), 'either text or text embedding must be supplied'
+        assert not (self.condition_on_text_encodings and (not exists(text_encodings) and not exists(text))), 'text encodings must be present if you specified you wish to condition on it on initialization'

-        times = torch.randint(0, self.num_timesteps, (b,), device = device, dtype = torch.long)
-        image_embed = self.get_image_embed(image)
-        text_cond = self.get_text_cond(text)
+        if exists(image):
+            image_embed = self.get_image_embed(image)
+
+        # calculate text conditionings, based on what is passed in
+
+        if exists(text):
+            text_cond = self.get_text_cond(text)
+        else:
+            text_cond = dict(
+                text_embed = text_embed,
+                text_encodings = text_encodings,
+                mask = text_mask
+            )
+
+        # timestep conditioning from ddpm
+
+        batch, device = image_embed.shape[0], image_embed.device
+        times = torch.randint(0, self.num_timesteps, (batch,), device = device, dtype = torch.long)
+
+        # calculate forward loss

        loss = self.p_losses(image_embed, times, text_cond = text_cond, *args, **kwargs)
        return loss
@@ -1116,12 +1158,13 @@ class Decoder(nn.Module):
        unet,
        *,
        clip,
-        vae = None,
+        vae = tuple(),
        timesteps = 1000,
        cond_drop_prob = 0.2,
        loss_type = 'l1',
        beta_schedule = 'cosine',
-        predict_x0 = False,
+        predict_x_start = False,
+        predict_x_start_for_latent_diffusion = False,
        image_sizes = None,                         # for cascading ddpm, image size at each stage
        lowres_cond_upsample_mode = 'bilinear',     # cascading ddpm - low resolution upsample mode
        lowres_downsample_first = True,             # cascading ddpm - resizes to lower resolution, then to next conditional resolution + blur
@@ -1172,7 +1215,7 @@ class Decoder(nn.Module):

        # predict x0 config

-        self.predict_x0 = cast_tuple(predict_x0, len(unets))
+        self.predict_x_start = cast_tuple(predict_x_start, len(unets)) if not predict_x_start_for_latent_diffusion else tuple(map(lambda t: isinstance(t, VQGanVAE), self.vaes))

        # cascading ddpm related stuff

@@ -1292,31 +1335,31 @@ class Decoder(nn.Module):
        posterior_log_variance_clipped = extract(self.posterior_log_variance_clipped, t, x_t.shape)
        return posterior_mean, posterior_variance, posterior_log_variance_clipped

-    def p_mean_variance(self, unet, x, t, image_embed, text_encodings = None, lowres_cond_img = None, clip_denoised = True, predict_x0 = False, cond_scale = 1.):
+    def p_mean_variance(self, unet, x, t, image_embed, text_encodings = None, lowres_cond_img = None, clip_denoised = True, predict_x_start = False, cond_scale = 1.):
        pred = unet.forward_with_cond_scale(x, t, image_embed = image_embed, text_encodings = text_encodings, cond_scale = cond_scale, lowres_cond_img = lowres_cond_img)

-        if predict_x0:
+        if predict_x_start:
            x_recon = pred
        else:
            x_recon = self.predict_start_from_noise(x, t = t, noise = pred)

-        if clip_denoised and not predict_x0:
+        if clip_denoised and not predict_x_start:
            x_recon.clamp_(-1., 1.)

        model_mean, posterior_variance, posterior_log_variance = self.q_posterior(x_start=x_recon, x_t=x, t=t)
        return model_mean, posterior_variance, posterior_log_variance

    @torch.no_grad()
-    def p_sample(self, unet, x, t, image_embed, text_encodings = None, cond_scale = 1., lowres_cond_img = None, predict_x0 = False, clip_denoised = True, repeat_noise = False):
+    def p_sample(self, unet, x, t, image_embed, text_encodings = None, cond_scale = 1., lowres_cond_img = None, predict_x_start = False, clip_denoised = True, repeat_noise = False):
        b, *_, device = *x.shape, x.device
-        model_mean, _, model_log_variance = self.p_mean_variance(unet, x = x, t = t, image_embed = image_embed, text_encodings = text_encodings, cond_scale = cond_scale, lowres_cond_img = lowres_cond_img, clip_denoised = clip_denoised, predict_x0 = predict_x0)
+        model_mean, _, model_log_variance = self.p_mean_variance(unet, x = x, t = t, image_embed = image_embed, text_encodings = text_encodings, cond_scale = cond_scale, lowres_cond_img = lowres_cond_img, clip_denoised = clip_denoised, predict_x_start = predict_x_start)
        noise = noise_like(x.shape, device, repeat_noise)
        # no noise when t == 0
        nonzero_mask = (1 - (t == 0).float()).reshape(b, *((1,) * (len(x.shape) - 1)))
        return model_mean + nonzero_mask * (0.5 * model_log_variance).exp() * noise

    @torch.no_grad()
-    def p_sample_loop(self, unet, shape, image_embed, predict_x0 = False, lowres_cond_img = None, text_encodings = None, cond_scale = 1):
+    def p_sample_loop(self, unet, shape, image_embed, predict_x_start = False, lowres_cond_img = None, text_encodings = None, cond_scale = 1):
        device = self.betas.device

        b = shape[0]
@@ -1331,7 +1374,7 @@ class Decoder(nn.Module):
                text_encodings = text_encodings,
                cond_scale = cond_scale,
                lowres_cond_img = lowres_cond_img,
-                predict_x0 = predict_x0
+                predict_x_start = predict_x_start
            )

        return img
@@ -1344,7 +1387,7 @@ class Decoder(nn.Module):
            extract(self.sqrt_one_minus_alphas_cumprod, t, x_start.shape) * noise
        )

-    def p_losses(self, unet, x_start, t, *, image_embed, lowres_cond_img = None, text_encodings = None, predict_x0 = False, noise = None):
+    def p_losses(self, unet, x_start, t, *, image_embed, lowres_cond_img = None, text_encodings = None, predict_x_start = False, noise = None):
        noise = default(noise, lambda: torch.randn_like(x_start))

        x_noisy = self.q_sample(x_start = x_start, t = t, noise = noise)
@@ -1358,7 +1401,7 @@ class Decoder(nn.Module):
            cond_drop_prob = self.cond_drop_prob
        )

-        target = noise if not predict_x0 else x_start
+        target = noise if not predict_x_start else x_start

        if self.loss_type == 'l1':
            loss = F.l1_loss(target, x_recon)
@@ -1380,7 +1423,7 @@ class Decoder(nn.Module):

        img = None

-        for unet, vae, channel, image_size, predict_x0 in tqdm(zip(self.unets, self.vaes, self.sample_channels, self.image_sizes, self.predict_x0)):
+        for unet, vae, channel, image_size, predict_x_start in tqdm(zip(self.unets, self.vaes, self.sample_channels, self.image_sizes, self.predict_x_start)):
            with self.one_unet_in_gpu(unet = unet):
                lowres_cond_img = None
                shape = (batch_size, channel, image_size, image_size)
@@ -1400,7 +1443,7 @@ class Decoder(nn.Module):
                    image_embed = image_embed,
                    text_encodings = text_encodings,
                    cond_scale = cond_scale,
-                    predict_x0 = predict_x0,
+                    predict_x_start = predict_x_start,
                    lowres_cond_img = lowres_cond_img
                )

@@ -1424,7 +1467,7 @@ class Decoder(nn.Module):

        target_image_size = self.image_sizes[unet_index]
        vae = self.vaes[unet_index]
-        predict_x0 = self.predict_x0[unet_index]
+        predict_x_start = self.predict_x_start[unet_index]

        b, c, h, w, device, = *image.shape, image.device

@@ -1448,7 +1491,7 @@ class Decoder(nn.Module):
            if exists(lowres_cond_img):
                lowres_cond_img = vae.encode(lowres_cond_img)

-        return self.p_losses(unet, image, times, image_embed = image_embed, text_encodings = text_encodings, lowres_cond_img = lowres_cond_img, predict_x0 = predict_x0)
+        return self.p_losses(unet, image, times, image_embed = image_embed, text_encodings = text_encodings, lowres_cond_img = lowres_cond_img, predict_x_start = predict_x_start)

 # main class

--- a/dalle2_pytorch/train.py
+++ b/dalle2_pytorch/train.py
@@ -0,0 +1,53 @@
+import copy
+import torch
+from torch import nn
+
+# exponential moving average wrapper
+
+class EMA(nn.Module):
+    def __init__(
+        self,
+        model,
+        beta = 0.99,
+        ema_update_after_step = 1000,
+        ema_update_every = 10,
+    ):
+        super().__init__()
+        self.beta = beta
+        self.online_model = model
+        self.ema_model = copy.deepcopy(model)
+
+        self.ema_update_after_step = ema_update_after_step # only start EMA after this step number, starting at 0
+        self.ema_update_every = ema_update_every
+
+        self.register_buffer('initted', torch.Tensor([False]))
+        self.register_buffer('step', torch.tensor([0.]))
+
+    def update(self):
+        self.step += 1
+
+        if self.step <= self.ema_update_after_step or (self.step % self.ema_update_every) != 0:
+            return
+
+        if not self.initted:
+            self.ema_model.state_dict(self.online_model.state_dict())
+            self.initted.data.copy_(torch.Tensor([True]))
+
+        self.update_moving_average(self.ema_model, self.online_model)
+
+    def update_moving_average(ma_model, current_model):
+        def calculate_ema(beta, old, new):
+            if not exists(old):
+                return new
+            return old * beta + (1 - beta) * new
+
+        for current_params, ma_params in zip(current_model.parameters(), ma_model.parameters()):
+            old_weight, up_weight = ma_params.data, current_params.data
+            ma_params.data = calculate_ema(self.beta, old_weight, up_weight)
+
+        for current_buffer, ma_buffer in zip(current_model.buffers(), ma_model.buffers()):
+            new_buffer_value = calculate_ema(self.beta, ma_buffer, current_buffer)
+            ma_buffer.copy_(new_buffer_value)
+
+    def __call__(self, *args, **kwargs):
+        return self.ema_model(*args, **kwargs)
--- a/dalle2_pytorch/vqgan_vae.py
+++ b/dalle2_pytorch/vqgan_vae.py
@@ -13,6 +13,8 @@ import torchvision

 from einops import rearrange, reduce, repeat

+from dalle2_pytorch.attention import QueryAttnUpsample
+
 # constants

 MList = nn.ModuleList
@@ -243,6 +245,7 @@ class ResBlock(nn.Module):
    def forward(self, x):
        return self.net(x) + x

+# vqgan attention layer
 class VQGanAttention(nn.Module):
    def __init__(
        self,
@@ -375,7 +378,7 @@ class VQGanVAE(nn.Module):

        for layer_index, (dim_in, dim_out), layer_num_resnet_blocks, layer_use_attn in zip(range(layers), dim_pairs, num_resnet_blocks, use_attn):
            append(self.encoders, nn.Sequential(nn.Conv2d(dim_in, dim_out, 4, stride = 2, padding = 1), leaky_relu()))
-            prepend(self.decoders, nn.Sequential(nn.Upsample(scale_factor = 2, mode = 'bilinear', align_corners = False), nn.Conv2d(dim_out, dim_in, 3, padding = 1), leaky_relu()))
+            prepend(self.decoders, nn.Sequential(nn.ConvTranspose2d(dim_out, dim_in, 4, 2, 1), leaky_relu()))

            if layer_use_attn:
                prepend(self.decoders, VQGanAttention(dim = dim_out, heads = attn_heads, dim_head = attn_dim_head, dropout = attn_dropout))
--- a/setup.py
+++ b/setup.py
@@ -10,7 +10,7 @@ setup(
      'dream = dalle2_pytorch.cli:dream'
    ],
  },
-  version = '0.0.41',
+  version = '0.0.48',
  license='MIT',
  description = 'DALL-E 2',
  author = 'Phil Wang',
Author	SHA1	Message	Date
Phil Wang	7ba6357c05	allow for training the Prior network with precomputed CLIP embeddings (or text encodings)	2022-04-26 09:29:51 -07:00
Phil Wang	76e063e8b7	refactor so that the causal transformer in the diffusion prior network can be conditioned without text encodings (for Laions parallel efforts, although it seems from the paper it is needed)	2022-04-26 09:00:11 -07:00
Phil Wang	4d25976f33	make sure non-latent diffusion still works	2022-04-26 08:36:00 -07:00
Phil Wang	0b28ee0d01	revert back to old upsampling, paper does not work	2022-04-26 07:39:04 -07:00
Phil Wang	45262a4bb7	bring in the exponential moving average wrapper, to get ready for training	2022-04-25 19:24:13 -07:00
Phil Wang	13a58a78c4	scratch off todo	2022-04-25 19:01:30 -07:00
Phil Wang	f75d49c781	start a file for all attention-related modules, use attention-based upsampling in the unets in dalle-2	2022-04-25 18:59:10 -07:00
Phil Wang	3b520dfa85	bring in attention-based upsampling to strengthen vqgan-vae, seems to work as advertised in initial experiments in GAN	2022-04-25 17:27:45 -07:00
Phil Wang	79198c6ae4	keep readme simple for reader	2022-04-25 17:21:45 -07:00
Phil Wang	77a246b1b9	todo	2022-04-25 08:48:28 -07:00
Phil Wang	f93a3f6ed8	reprioritize	2022-04-25 08:44:27 -07:00
Phil Wang	8f2a0c7e00	better naming	2022-04-25 07:44:33 -07:00
Phil Wang	863f4ef243	just take care of the logic for setting all latent diffusion to predict x0, if needed	2022-04-24 10:06:42 -07:00