AI Images Without Paying A cent | Stable Diffusion Tutorial

 



Generative AI has opened the door to allowing anyone to create incredible images just using a computer and a few text prompts. Now, the platform that gets the least amount of attention is Stable Diffusion, but it might just be the hidden gem you're looking for.
Unlike other major generative AI platforms like Adobe's Firefly or OpenAI's Dali or even Midjourney, Stable Diffusion is completely open source, which in short is free. So, I'd want Padme to wear her iconic [music] white turtle attire. See that image is coming in. It's looking great. There's a few issues going on in the lower parts of the frame. In this particular section, I'm changing how the body structure looks. Again, let's increase that to four images. Okay, we're going to pause it right there. Remember, there are no limitations in the software, which means the responsibility is on you.
Stable diffusion lets you run on your own GPU instead of on the cloud, which allows you to tweak every minute setting that you can think of. Now, if that's not enough to convince you to learn this software, it's also the only platform with zero limitations, but it also comes with its own drawbacks. It's tough to get started with, but that's where we come in. Today, we're going to help you with everything from getting started, installation to even more complicated things like perfecting your prompts and getting the best possible image. Now, today's episode is a little bit different. It's just you, me, and a computer. And we're going to get through this together. It's also the first time that we haven't had someone behind the camera, so expect things to go wrong.
First off, what is stable diffusion? In short, it's an image generation model created by stable AI. It turns text  


prompts into images. What makes stable diffusion really special is that it's open source, which in short means that it's free, but it also means that you can download it, bring it into your own computer, and customize it to your own art style.

Now, in a majority of this video, we will work with a software application called Fucus. This is a front-end interface that allows us to interact with the stable diffusion model. Now, don't let any of those words scare you off. Think of stable diffusion as the engine, the code that's doing all the AI image processing. Focus is like the car that's built around the engine. It's what has all the buttons and the controls and the interface that allows you to interact with the engine without really messing around with the tech.

Okay, enough talking. Let's jump into the application. The first thing you want to do is go to the GitHub repository for Focus. We'll leave a link in the description. After you download the file, you want to extract the file and you'll see and run.bat. You want to click on the run.bat file and start the installation process. Remember, the download file will be around 50 GB. So, make sure you have that capacity before you get started. By default, you'll get the standard model, but you'll also get the option to run the animate and realistic models as well. For right now, we're not going to change any of the default behavior of the software, which means as it launches, it'll look for new models, keep the software updated, which is generally what you want. But if you want to change that, you can use these two command lines. Let's do realistic right now. And now we just let it do its thing. Another quick tip that I like to use is keeping task manager open. As long as I see activity on the GPU, I know that the software is working with the GPU to get it collected.

Awesome. And we are finally ready to start generating images. You'll notice a few things. The image generation window at the top, the text prompt window at the bottom where we can type things like cats on a window lid. And at the bottom, you can see input image. We'll go into that in detail. It's something that I use in great detail. Pain enhance, which we won't talk about too much.

essentially post-processing steps that you can take on your final image to increase the resolution and detail. As we hit generate, you can see that there is an immediate spike in your GPU. So, you'll notice the image comes in a little fuzzy, then it cleans up over time. Okay, with those images in place, we can look at them in detail by clicking on each individual image. The second image is a really good example of the kind of problems that this kind of model has. You can see that the eye on the left is perfect, but the eye on the right just doesn't have the detail we need. So, there's two things that we can think about at this point. We can click on advanced and we can first look at the presets. We are using the realistic model which is what generates such fine detail but there's lots of other variations that you can try out here.

There's also performance depending on if you want to take a qualitative or quantitative approach. Right now we're prioritizing speed of delivery, but we can also change that to quality to generate higher quality images. Now a feature that I use all the time is changing the number of images, but for right now let's do four images.

>> [music] >> That's pretty incredible. The first three images have come in. I love the detail here. There's lots of detail in the hair and in the eyes. Both eyes look perfect. The second image is not quite as good. It has really great perspective. Lots of detail in the bricks and the window, but not quite as much detail on the cat itself. You can see that the eye on the right isn't perfect. Okay, this is great perspective cuz there's a beautiful window. It's a low angle shot. Again, not quite as much detail on the cat itself. And the final image has also come through. Oh, this is amazing. And you can see the cat looking through the window. Lots of detailing on the glass and the light and the texture on the skin. We're going to build a crazy Star Wars poster no one has ever seen of Padme fighting off Anakin Skywalker. But first, let's build a cyberpunk city with crazy [music] detail.

So, you can notice the kind of detail that I'm putting in here. I'm I'm writing out very specifically what I want to see. So, a cyberpunk city at night. So, I'm specifying what the lighting looks like. Glowing neon lights. someone have that look where there's lots of light from the buildings themselves. And I'm specifically calling for cinematic lighting. With AI models, the more specific details you can give, the more likely you are to get a successful final image. Okay, so those two images have come through. They're both fairly similar, which is something I don't really like. Ideally, I want variation on shot. Also, something I like to do is watch the model as the image generates just to see if it's something that's in alignment with what I want. If it's completely off, I can choose to skip through that image. Okay.

And that's pretty great. Everything's generally looking okay, but nothing is specifically looking good. That's also another problem with the AI generation model for wide and detailed shots. It's harder to get fine details correct when there's lots of details spread across the image. Let's say we love this image in general, but we want to enhance certain details. With that in mind, let's refine our model some more. Okay, now we've got two new images. Let's look at both of them. I love the detailing on this. Generally, everything seems okay.

There's lots of lost details in the buildings in the background, but I think I can generally fix those as we go. The second image is a lot softer and there is a lot of billboards and written detail which [music] I probably won't be able to fix. What if we have a great image, but we want to refine that image.

That's where image input comes in.

First, if you want to take this image and we just want to scale it up, bring in new resolution, we can use image upscale. So, you just drag that image in here. First, we look at the bottom row.

Upscaling by 1.5, upscaling by 2x. that essentially just expands the resolution, maintaining as much of the image as possible. And then you have upscaling fast 2x, which is the same thing, but it processes it a little faster with a little less accuracy. So, first let's do a 1.5 upscale. Now, as that image comes in, we can see we already have a lot more detail. Okay, the first image is in. Let's take a look at what that looks like. So, you're already seeing so much more detail in the building in the foreground. All the problems of the background have now been resolved. All the lines are straight. All the windows are visible. Generally, everything's looking good. There is still a little bit of specific problems I'm seeing, especially on this billboard on the right, parking lot down below, but the cars don't look perfect. We can work on those specifically. Okay, let's look at this image with some detail. That's not going to work cuz it's so close to the foreground element, which is this building. So, to refine this, we're going to use impaint. This allows you to refine and work on specific details of your image with great detail. There's a brush, and anything you brush over will then be changed. Anything that's outside the brush radius will not be affected.

So, we've got three options here.

Impaint, improve, and modify. Impaint is a great technique if you want to change something of your image with subtle variation and keep the general look of the image the same. The first thing you'll notice is that the GPU is using all of its processing power on just that one zone of the image, which means you'll get a lot of resolution in just that one area. Has a little motel looking building with a swimming pool in front. It's lots of details in terms of cars and people, which may not be what we want because it's going to attract a lot of attention. And then the second is a black building with a few windows. You can see that I can actually drag the image from my image generation window into my impaint window. This is a pretty common technique of refining your image as you move it back and forth within the software itself. And this is a good time to talk about the other two impend features as well. Improve detail is used very often. This is when you want to increase the resolution of something in the background, right? We can see far more detail in that image. You can see the specific floors of the building. You can see through some of the windows. You can see some trees and shbery brought in front of the building. Really quickly, let's look at modify content as well.

This is a powerful tool when you want to make a dramatic shift of your image and then later go back and refine it using the refineer tool. So, let's look at one of these. And that's what that car path looks like. Again, here you can see a lot of imperfections in the car, in the floor, in the building next to the car park. All of which will need to be refined if you want to use it in your final image. But for right now, we're just going to revert back to the image that we had and look through our final settings. At this point, you probably get the idea of image generation, but let's dive into something a little bit more obscure that uses more advanced features of the software.

So, [music] let's think of something that we can't find on the internet.

though. Anakin Skywalker and Padme lightsaber dual Star Wars franchise high detail cinematic lighting at night.

[music] All right, let's see what this looks like. Again, let's increase that to four images. Okay, we're going to pause it right there. So, this is where you need to be careful as you generate images. You need to make sure that you don't generate any images that's not safe for work. And this is a good time as any to talk about the power of these creative tools. Remember, there are no limitations in these software, which means the responsibility is on you. Make sure you don't create any content that's harmful, misleading, or violates privacy. Another tool that we can use to safeguard our content is negative prompts. Now, these are things that you don't want in your image. So, by default, I'm getting things like unrealistic, saturated, big nose, painting, drawing, sketch. These are all prompts that have come in default by the software. But now, I'm going to also include not safe for work as a prompt.

And you can expand on that list. Okay.

And those images have come through. They both have their own challenges. So over here, both Anakin and Padme are sharing a lightsaber, which doesn't really make the most sense. In the second image, I kind of have their hands crossed over.

Both which we can fix. And we can do that by bringing this image, dragging it in using impaint and correcting for. But we're not going to do that right now because there is a holistic problem in the image. And that's the fact that I don't like the perspective that we're getting. Ideally, I'd want to see something similar to the Star Wars poster. It's a low angle shot with lava in the background and each character fighting aggressively against each other. Now I could give this in the form of a prompt, but I can also use image prompt. Image prompt essentially allows you to input an image and generate a similar image. Now, but for right now, let's turn off all of our advanced features. You can go over to image prompt and you can drag your image in and you have a few features that will show up. Now, if you don't have this bar at the bottom, you just want to scroll down to the bottom and click advanced.

That'll give you four options. Now, image prompt will generally scan the image very similar to a textbased prompt. It'll attempt to understand what's happening in the image and it'll generally use that as a suggestion for generation. You can see that it's generally taking the idea of the image and generating something similar. But in this case, what we really need is something that's visually similar. So, we have two real options. Pyrochni and CPDS. Pyrochani essentially maps the position of characters and takes it into your new image which would be ideal for this particular image and CPDS essentially uses contrast, color and saturation to generate a similar image.

So in this particular case, Pyrocheni would work perfect. Okay, we've got four images in. Let's let the remaining images come in. Let's look at the first one. So this first image, the first thing you'll notice is that it's fairly low resolution. There's not a lot of detail in the face structure. The lightsaber's broken. I don't like how the sand looks relative to the mountain.

The second image is much better. And then it's converted Obi-Wan into some version of Anakin. So, that's really great in terms of perspective, in terms of each lightsaber. Lightsabers are red, but I'm assuming I can make those changes. Third image. I don't like it as much. The resolution is not quite there.

The perspective isn't great. Anakin's looking a lot bigger than Padme. It's probably one we're going to avoid. Image four. That's looking much better. I love the perspective of this. This next image is completely broken. perspectives are wrong. Seems to be a dual lightsaber which Padme is holding from the wrong side. So, we're not going to use that image. Final image looks like two Padme fighting each other. Again, it's less of a problem because I know I can change one of those characters, [music] but I'm not a big fan of the perspective. I'm not a big fan of the background. So, with all of that in mind, I think I'm going to go with this image as the image I'm going to refine to final.

Okay. Okay. So, the first thing I'll do is drag that image into impaint and start erase my previous painting and start working on specific sections of the image. So, I'd want Padme to wear her iconic white battle attire. So, these images are already more in direction of where it needs to be. This time, I'm only going to change up the outfit itself. Just highlighting very specifically. So, our first section of images are coming in and they're all looking pretty great. So, I love the position here. Doesn't look perfect in terms of outfit. Second outfit looks much better, but there's lots of problems with the hand structure. The hand seems a lot smaller than it needs to be. The foot structure isn't great.

This third image is better in terms of hand structure, but there is a broken limb right here, and there's a few issues going on in the lower parts of the frame. So, here's a pro tip. In this particular section, I'm changing how the body structure looks, but I don't really want to change the general position. I find that it's helpful to highlight only parts of the image. So you'll notice here I've left the foot open and the back of the ankle open as well along with the waist of the character. That way I can change just the middle section of the image without affecting the overall look of the Now I'm liking the overall structure here. I just want to fix this broken limb. I'd like to change the shoes the character is wearing.

Maybe reduce some of this armor. And that's what I'm going to do now. Now you can see that after a certain point I'm going to start running parallel processes. In this section I've highlighted Padme's hair very specifically. So I'm trying to get her hair tied up and it's proving to be a little bit difficult. But the important point here is that you need to keep a reference ready. I'm trying to match those references. Having a reference will always help get the details right and anchor your character in reality.

So, one thing we'll definitely need to do is work on character features, especially in the face. You want to have as much detail as possible in the face.

So, even if things aren't perfect, most people won't notice. And you can see I'm already painting in the second image while that first image is generated. I'm going to need to work on that next. So, Anakin, Star Wars, detailed face, it's a lot better. We're not going to perfect anything. We're going to do the best we can and really just go through all of the different tools and functions. If you'd want to see us explore that in another video where we deep dive into creating hyper realistic images, let us know in the comments below. Now, in this next section, I want to show you how to expand an image. Now, you've probably seen something similar if you've used Adobe's Photoshop with content of airfield, but it's quite powerful in focus because you can really control and refine how the expansion work prompt you put in here is very specifically the information that you want in the background and in the expanded areas.

But ideally, I'd like to have this in landscape format. So, what I'm going to do is I'm going to take that image, put it back into content expansion, and just generate the left and right side of the image. Okay, that's looking really great. I love how that looks. Now, I don't really like how the characters are standing. I'm seeing a lot of imperfections in the hand structure, mannequin's leg structure, but generally this is a great starting base considering we generated this image in under an hour. That's a really great starting point. Now, there's a lot more you can do here using the refiner tools.

You can bring in things like smoke and fog. You can work on the character outfits, bring in detail, so things like the shoes and the hair. I'll also look back at my references and really understand where I missed the ball. And that brings us to the end of this episode. If you like this video, hit the like and subscribe buttons. Now, there is a lot more we can do to perfect this photo.

Comments