Vibrational spectroscopy · research platform

Reading the fingerprint region — at research depth.

Chemdu IR is a working environment for infrared and vibrational spectroscopy: a referenced band atlas with explicit normal-mode assignments, biomolecular IR, advanced experimental methods, computational spectral prediction, and an AI program that mines the information-dense fingerprint region for molecular identity, noncovalent interactions, and protein structure.

400–4000 cm⁻¹ working range GFN2-xTB spectra on demand <1500 cm⁻¹ — the target region

The research program

Why the fingerprint region

Above ~1500 cm⁻¹, IR bands are local: a carbonyl, a nitrile, an N–H. They report functional class. Below 1500 cm⁻¹ the spectrum becomes a dense manifold of coupled skeletal modes whose pattern is specific to a single molecule and exquisitely sensitive to conformation, environment, and noncovalent contacts. That sensitivity is exactly what makes the region hard to read by eye — and what makes it a target for machine learning. The platform is organised around four objectives:

Objective 1

Compound identification

Map a measured or computed spectrum to molecular identity using the fingerprint pattern rather than a handful of group frequencies. Most tractable of the four.

Objective 2

Noncovalent interactions

Detect and quantify H-bonding, halogen bonding and ion pairing from donor-mode red-shifts, broadening and intensity gains (Badger–Bauer regime).

Objective 3

Secondary structure

Resolve protein/peptide secondary structure from amide I band shape — α-helix, β-sheet, turns — with isotope editing and 2D-IR for residue resolution.

Objective 4

Protein–ligand binding

Read binding-pocket electrostatics and ligand contacts via vibrational Stark probes and reaction-induced difference spectra. The frontier objective.

Scope — what IR can and cannot resolve

IR does not read a primary sequence (that is mass-spectrometry territory). It reports composition and secondary structure, not residue order. Compound identification is the most achievable objective; noncovalent and protein–ligand readouts are genuine research frontiers that depend on vibrational-probe design, difference spectroscopy, and 2D-IR. This platform is explicit about that boundary throughout.

Modules

Work areas

Reference

Band atlas

Group frequencies with explicit normal-mode and symmetry assignments, intensities, a searchable table, and an interactive peak identifier.

open →
Data · experimental

Experimental IR database

A searchable, digitized library of real gas-phase FTIR spectra (NIST WebBook) across the major functional classes, with an on-site viewer.

open →
Data

Amino-acid IR

Referenced bands for all 20 side chains (Barth compilation) — carboxyl, guanidinium, ammonium, aromatic and amide markers — with a simulator.

open →
Data

Nucleic-acid IR

DNA/RNA base carbonyl & ring markers, sugar-phosphate backbone, and B / A / Z / RNA conformational marker bands.

open →
Biomolecular

Protein & peptide IR

Amide A/I/II/III modes, a secondary-structure band map, an amide-I composition estimator, H/D exchange and isotope-edited 2D-IR.

open →
Simulate · on-site

IR spectrum simulator

Overlay real NIST measured spectra and reconstructed band profiles of amino acids, nucleotides, conformers and secondary structures — in-browser, no external service.

open →
Methods

Advanced techniques

ATR / DRIFTS / PM-IRRAS sampling, 2D-IR, time-resolved & difference IR, VCD, SFG, and vibrational Stark probes.

open →
Computation

Predicting spectra

Harmonic vs anharmonic frequencies, scaling factors, VPT2, the GFN2-xTB pipeline, and generating synthetic training data at scale.

open →
Objective 3

IR sequencing

Can IR read an amino-acid sequence? An honest, interactive study of residue-resolved IR and the reduced alphabet IR can actually resolve.

open →
AI program

Fingerprint & ML

The core research line — including a Phase-1 proof-of-concept: spectrum→functional-group and fingerprint-vs-high-region compound retrieval.

open →

Methodology

The compute keystone

The bottleneck for machine learning on vibrational spectra is labelled data. Measured reference libraries are finite and inconsistently conditioned; the platform supplements them with computed spectra generated on demand. Quantum-mechanical IR (GFN2-xTB Hessian + double-harmonic intensities) yields, for any structure, a peak list that is broadened to a fixed-length vector and labelled automatically by substructure pattern matching — producing arbitrarily large, perfectly-annotated training sets.

spectrum(structure) → GFN2-xTB Hessian → peak list (νi, Ii) → Lorentzian broaden → 181-pt vector [400–4000, 20 cm⁻¹] → label by SMARTS

Phase 1 of this pipeline is already validated on a curated 146-molecule set; results, methods and the honest caveats are on the fingerprint program page.

Foundations

Selected references

  1. Socrates, G. Infrared and Raman Characteristic Group Frequencies, 3rd ed. (Wiley) — group-frequency assignments.
  2. Krimm, S. & Bandekar, J. Vibrational spectroscopy and conformation of peptides, polypeptides, and proteins, Adv. Protein Chem. (1986) — amide-mode theory.
  3. Barth, A. Infrared spectroscopy of proteins, Biochim. Biophys. Acta (2007) — amide I secondary-structure analysis.
  4. Bannwarth, Ehlert & Grimme. GFN2-xTB, J. Chem. Theory Comput. (2019) — the semi-empirical method used for on-demand spectra.
  5. NIST Chemistry WebBook & CCCBDB — reference gas-phase IR and computed-frequency scaling factors.

Internal research resource — not for instruction. Group-frequency ranges are representative; exact positions depend on phase, conjugation, ring strain, hydrogen bonding and environment.