Research objectGeneralization bound
Out-of-domain generalization from rephrasability and stability
Rephrasability of the data-generating process and stability of the predictor class jointly control the gap between in-domain denoising and prediction on unseen domains. For unbounded autoregressive processes, this gives an out-of-domain generalization guarantee at a token-level rate.
Statement
\[\sup_{\theta\in\Theta_K}\left|L_{\mathrm{out}}(\theta)-L_{\mathrm{in}}(\theta)\right|\leq\widetilde{\mathcal O}\!\left(\rho\sqrt{\frac{C(\varepsilon)}{NT}}\right)\]
- \(N\)
- number of independent sequences
- \(T\)
- number of predicted tokens per sequence
- \(C(\varepsilon)\)
- sequential metric entropy at resolution \(\varepsilon\)
- \(\rho\)
- effective dependence scale determined by the rephrasability condition
- \(\Theta_K\)
- stable hypothesis class
- \(\widetilde{\mathcal O}\)
- asymptotic upper-bound notation suppressing logarithmic factors in the other problem parameters and confidence level
- \(L_{\mathrm{in}}\)
- in-sample denoising error
- \(L_{\mathrm{out}}\)
- out-of-domain prediction error