- Home
- Documentation
- API reference
- Spectral
- hyperproc.spectral._smoothspline
hyperproc.spectral._smoothspline¶
Source-derived reference
Generated from the current hyperproc 0.1.2 checkout.
Implementation: hyperproc/spectral/_smoothspline.py. Signatures, defaults, docstrings, and expandable source are extracted statically; the module is not imported or executed. Names beginning with _ are implementation details, not a stable public API.
Use the function signature as the authority for individual parameter defaults and return annotations. Original docstrings sometimes group parameter names or wrap return descriptions across lines; these descriptions are preserved rather than inferred or rewritten.
Exported entry points¶
| Importable name | Definition |
|---|---|
hyperproc.spectral._smoothspline.SmoothSpline |
hyperproc.spectral._smoothspline.SmoothSpline |
hyperproc.spectral._smoothspline.smooth_spline |
hyperproc.spectral._smoothspline.smooth_spline |
hyperproc.spectral._smoothspline.nknots_smspl |
hyperproc.spectral._smoothspline.nknots_smspl |
A faithful port of R's stats::smooth.spline, for reproducing R pipelines.
Published spectroscopy workflows are often written in R, and a Python package
that wants to reproduce their numbers has to reproduce their smoother, not
merely a smoother. scipy's UnivariateSpline is a different algorithm
with a different parameterisation, so it cannot: it would give plausible
answers that do not match.
This follows R's implementation step for step, from src/library/stats:
.nknots.smsplchooses the number of interior knots (smspline.R).xis de-meaned, rounded to a tolerance of1e-6 * IQR(x)and de-duplicated, then scaled to the unit interval.- knots are taken at
xbar[trunc(seq(1, nx, length.out = nknots))]with the end knots repeated three times, givingnknots + 2cubic B-splines. - the penalty is the Gram matrix of the basis second derivatives, computed exactly: a cubic B-spline's second derivative is piecewise linear, so the product is piecewise quadratic and three-point Gauss-Legendre is exact.
- the smoothing parameter enters as
lambda = ratio * 16^(6*spar - 2)withratio = tr(X'WX) / tr(Sigma)over the interior band (sbart.c). dfis matched by minimising3 + (df - tr(H))^2oversparin[-1.5, 1.5], using Brent's golden-section and parabolic search with R's own constants (tol = 1e-4,eps = 2e-8, 500 iterations).
That last point is why the search is transcribed rather than replaced by a
root find. R stops when spar is known to about 1e-4, so its df lands
near 59.995 rather than 60. An exact solver would be more accurate and would
not match. Reproducing the published number means reproducing the search.
_GOLD = 0.38196601125010515
module-attribute
¶
_BIG = 1e+100
module-attribute
¶
SmoothSpline
¶
The fitted spline, with R's own diagnostics on it.
Attributes mirror R's object: df (the achieved trace, not the target),
lambda_, spar, nknots and coef.
Source code in hyperproc/spectral/_smoothspline.py
156 157 158 159 160 161 162 163 164 165 166 167 168 169 170 171 172 173 174 175 176 177 178 179 180 181 182 183 184 185 186 187 188 189 190 191 192 193 194 195 196 197 198 199 200 201 202 203 204 205 206 207 208 209 210 211 212 213 214 215 216 217 218 219 220 221 222 223 224 225 226 227 228 229 230 231 232 233 234 235 236 237 238 239 | |
knot = _knot_vector(xbar, nknots)
instance-attribute
¶
nknots = nknots
instance-attribute
¶
ratio = ratio
instance-attribute
¶
spar = np.nan
instance-attribute
¶
lambda_ = float(lam)
instance-attribute
¶
spar_used = last
instance-attribute
¶
lev = lev
instance-attribute
¶
df = float(lev.sum())
instance-attribute
¶
__init__(x, y, df=None, spar=None, lam=None, w=None, tol=None, nknots=None, maxit=500, search_tol=0.0001, eps=2e-08)
¶
Source code in hyperproc/spectral/_smoothspline.py
predict(x) -> np.ndarray
¶
Evaluate the spline, extrapolating linearly as a natural spline does.
Source code in hyperproc/spectral/_smoothspline.py
nknots_smspl(n: int) -> int
¶
R's .nknots.smspl: how many interior knots for n unique points.
Source code in hyperproc/spectral/_smoothspline.py
_iqr(x: np.ndarray) -> float
¶
R's IQR, which uses quantile type 7 (numpy's default).
_collapse(x, y, w, tol)
¶
De-duplicate x to a tolerance, averaging y as R does, and sort.
Source code in hyperproc/spectral/_smoothspline.py
_knot_vector(xbar: np.ndarray, nknots: int) -> np.ndarray
¶
c(rep(xbar[1],3), xbar[seq(1, nx, length.out=nknots)], rep(xbar[nx],3)).
R indexes with the raw doubles from seq.int, and indexing with a double
truncates, so the interior knots are at truncated positions, not rounded.
Source code in hyperproc/spectral/_smoothspline.py
_basis(xbar: np.ndarray, knot: np.ndarray) -> np.ndarray
¶
Source code in hyperproc/spectral/_smoothspline.py
_gram(knot: np.ndarray) -> np.ndarray
¶
Sigma[i, j] = integral B''_i B''_j, exactly as R's sgram.f computes it.
A cubic B-spline's second derivative is piecewise linear, so the integral of
a product over one knot interval has the closed form
w * (a*c + (a*d + b*c)/2 + b*d/3). R writes that last coefficient as the
literal 0.3330, not 1/3, and has done since the original Fortran.
That looks like a rounding of no consequence and is not: it shifts the
penalty by about 1e-3 relative, which moves ratio = tr(X'WX)/tr(Sigma)
by the same amount and therefore moves lambda. Integrating exactly, by
Gauss-Legendre, is the more correct thing to do and disagrees with R in the
third decimal of the fitted spectrum. Reproducing R means reproducing this.
Source code in hyperproc/spectral/_smoothspline.py
_solve(X, sigma, wbar, ybar, lam)
¶
Fit at one lambda; return coefficients, fitted values and leverages.
Source code in hyperproc/spectral/_smoothspline.py
_brent_fmin(f, ax, bx, tol, eps, maxit)
¶
Brent's golden-section plus parabolic search, as transcribed in sbart.c.
Returns (best, last_evaluated). Both are needed because R reports the
first and fits at the second.
Not replaced by a root find on purpose. R stops once spar is known to
about tol, so its answer carries that error; solving exactly would be
more accurate and would not reproduce the published numbers.