Paper: Transformers learn in-context by gradient descent — AI Alignment Forum