Project
RamanUQ: Uncertainty Quantification for Raman Defect Metrics in Carbon Materials
An open, pre-registered pipeline showing standard I_D/I_G error bars are badly miscalibrated, plus a protocol resolving defect changes 1.2 to 2.8 times finer.
Motivation
Raman conclusions about defected carbon often rest on the I_D/I_G intensity ratio, which is computed only after a chain of quiet analysis choices: despiking, baseline correction, lineshape selection among pseudo-Voigt and Breit-Wigner-Fano forms, peak set, and intensity convention (height versus area). My own graphene oxide analysis showed how easily those choices could move a scientific claim. RamanUQ exists to make the choices visible and propagate them into honest uncertainty on the ratio.
Methods
The pipeline covers despiking, asymmetric-least-squares baseline correction, multi-peak fitting with pseudo-Voigt and BWF lineshapes, model selection with information criteria (AIC/BIC), and residual-bootstrap uncertainty, evaluated across a full configuration grid on synthetic spectra with known truth at multiple signal-to-noise ratios, plus digitized published spectra spanning ion-bombarded graphene, N-doped rGO, and functionalized CNTs. Calibration anchors come from Tuinstra and Koenig (1970) and Cançado et al. (2006, 2011). Four research questions (Q1, Q1b, Q2, Q3) and a dated Q2 prediction were pre-registered before any results were computed. Validation ran through explicit gates in continuous integration, including reproduction of published I_D/I_G values under both intensity conventions (height: 1.523 against a published 1.6; area: 1.714 against a published 1.64) and a differential gate in which every science-critical formula was independently re-implemented by a clean-room agent and asserted equal on hundreds of randomized inputs. I designed the specifications, acceptance tests, gates, and pre-registration; AI coding agents implemented to those specifications under separated implementer, reference, and reviewer roles, with the division of labor disclosed in the release notes.
Results
Three findings, all frozen at v0.1.0 and regenerated by the pipeline itself. First, no configuration earned a recommendation: zero cells in the configuration-by-SNR grid reached the pre-registered 0.90 empirical-coverage floor (the best observed was 0.80), so the pre-registered ranking is honestly empty. Second, standard reported 95 percent error bars on I_D/I_G contain the truth only 27.6, 24.0, and 18.3 percent of the time at SNR 15, 50, and 200. Third, the released protocol’s minimum detectable change beats a naive pipeline in all three SNR regimes; the naive approach needs 1.2 to 2.8 times larger defect changes to detect the same effect.
Current status
Released and citable. v0.1.0 is the release of record for the pre-registered results; v0.2.0 added a second reproduced-literature case and polish without touching any frozen claim. The package lives on GitHub (AvinGupta-ship-it/ramanuq) with Zenodo DOI 10.5281/zenodo.20918643, and a single script reproduces every figure and reported number from a fresh clone.
Future work
Adoption by a lab other than mine is the next milestone; the protocol cards, disclosure checklist, and minimum-detectable-change curves are built for working spectroscopists. The tool also feeds back into my experimental work by defining how large a defect-density change must be before a Raman claim about hydrogenation is defensible.