Language Models are Few-Shot Learners
Source: https://arxiv.org/abs/2005.14165
GPT-3 demonstrated that scaling language models improves task-agnostic few-shot performance. The 175B parameter model achieved strong results without fine-tuning.
Source: https://arxiv.org/abs/2005.14165
GPT-3 demonstrated that scaling language models improves task-agnostic few-shot performance. The 175B parameter model achieved strong results without fine-tuning.