Channel BannerChannel Banner
YBlNQK0Ao6gYBlNQK0Ao6g

Image GPT: Generative Pretraining from Pixels (Paper Explained)

30,728
Socialcounts.Org
Enter Fullscreen (f)
963
likes
75
comments
daily view 0
monthly view 0

Share

Live Analytics

Comments 0

Image GPT: Generative Pretraining from Pixels (Paper Explained) Analytics Table

Income Estimates for Image GPT: Generative Pretraining from Pixels (Paper Explained)

Based on this YouTube video's total view count of 30.7K views and industry-standard rates, the estimated total earning is $22 - $61 through ad revenue. Historical data is not yet available to calculate daily, weekly, or monthly averages.

About Image GPT: Generative Pretraining from Pixels (Paper Explained)

Explore Image GPT: Generative Pretraining from Pixels (Paper Explained) with 30,728 views, 963 likes, and 75 comments. Experience the impact of this video content that has captured audience attention.

BERT and GPT-2/3 have shown the enormous power of using generative models as pre-training for classification tasks. However, for images, pre-training is usually done with supervised or self-supervised objectives. This paper investigates how far you can get when applying the principles from the world of NLP to the world of images. OUTLINE: 0:00 - Intro & Overview 2:50 - Generative Models for Pretraining 4:50 - Pretraining for Visual Tasks 7:40 - Model Architecture 15:15 - Linear Probe Experiments 24:15 - Fine-Tuning Experiments 30:25 - Conclusion & Comments Paper: https://cdn.openai.com/papers/Generative_Pretraining_from_Pixels_V2.pdf Blog: https://openai.com/blog/image-gpt/ Code: https://github.com/openai/image-gpt Abstract: Inspired by progress in unsupervised representation learning for natural language, we examine whether similar models can learn useful representations for images. We train a sequence Transformer to auto-regressively predict pixels, without incorporating knowledge of the 2D input structure. Despite training on low-resolution ImageNet without labels, we find that a GPT-2 scale model learns strong image representations as measured by linear probing, fine-tuning, and low-data classification. On CIFAR-10, we achieve 96.3% accuracy with a linear probe, outperforming a supervised Wide ResNet, and 99.0% accuracy with full finetuning, matching the top supervised pre-trained models. An even larger model trained on a mixture of ImageNet and web images is competitive with self-supervised benchmarks on ImageNet, achieving 72.0% top-1 accuracy on a linear probe of our features. Authors: Mark Chen, Alec Radford, Rewon Child, Jeff Wu, Heewoo Jun, Prafulla Dhariwal, David Luan, Ilya Sutskever Links: YouTube: https://www.youtube.com/c/yannickilcher Twitter: https://twitter.com/ykilcher Discord: https://discord.gg/4H8xxDF BitChute: https://www.bitchute.com/channel/yannic-kilcher Minds: https://www.minds.com/ykilcher

About YouTube Real-Time View Count

With SocialCounts.org’s view counter, track your YouTube video’s live view count and YouTube likes count in real time with fast, reliable updates.

Watch every YouTube video live view count rise with our real-time YouTube views tracker—built for accuracy and minimal delay.

Follow YouTube real time views as they happen, using our dedicated view counter for YouTube videos.

Get up-to-date live view count on YouTube and see real-time growth with SocialCounts.org’s smart tracking tools.

Embed Widget

Parameters:

  • fullscreen=true - Fullscreen counter
  • graph=true - Live graph chart
  • counter=0/1/2 - Select counter (0=likes, 1=views, 2=comments)
URL

Click to copy the embed URL to your clipboard