Representation recovery by first-order ANIL
In the linear infinite-task setting and under the stated initialization and step-size conditions, first-order ANIL asymptotically removes the component orthogonal to the true shared subspace, makes the learned initialization represent the zero function, and recovers a non-degenerate representation of the shared subspace despite overparameterization. An appendix refinement further gives a \(1/t\) decay rate for the squared orthogonal component.
- \(t\)
- outer-loop training iteration
- \(\boldsymbol B_t\)
- learned representation matrix at iteration \(t\)
- \(\boldsymbol w_t\)
- shared head initialization at iteration \(t\)
- \(\boldsymbol B_\star\)
- orthonormal basis for the true shared representation subspace
- \(\boldsymbol B_{\star,\perp}\)
- orthonormal basis for the complement of the true representation subspace
- \(\boldsymbol\Lambda_\star\)
- positive-definite limiting covariance on the true shared subspace