hmmm…
my AI imaging experice took place with a graphics card with only 8GB of VRAM.
the important componet of an operation like this is not to throway your ideas, but to simplify your targets.
give yourself 1 thing at a time to adjust and improve on.
that alows you to take away winds and losses from the process. what you learn from target one can then inform target 2. and then target 3, etc.
so as a basica stratigy. start with still images, and work on how you describe images, and then grade how well the model is following the image description. why? becase there are atleast 2 different possible failiers here:
its possible all of the things your downloaded could help EVENTUALLY. but atm you have so many working componets in your workflow that its hard to judge what is imediately usefull information, and what information cant be deciphered yet. and because of that, its almost impossible for anyone to give instructions on your direct setup on how to fix it, because knowing the componets isnt the same as understanding the exact failiers.