Skip to content
Sign in

/lenses

Everything we can see.

93 measurements, each one stating what it measures, how to read it, and why it is worth measuring. A lens that cannot answer the third question does not ship.

01 Pick what you are here for

Nobody needs all 93 at once. Each profile is a working set: choose one to see only the lenses that answer its question.

02 File facts

What the container says, and what it cannot tell you.

Duration
file.duration
Voice & podcastMix & masterProvenanceDataset QAEnvironmentBroadcast & film post

Measures Length of the audio.

Reading it Seconds.

Why Sets what else can be claimed: below three seconds most verdicts are guesses.

Sample rate
file.sampleRate
ProvenanceDataset QABroadcast & film post

Measures Samples per second.

Reading it 44.1k and 48k are standard; 8k or 16k means it came through a phone or a codec.

Why Caps the highest frequency that can exist at all, so it bounds every spectral reading.

Channels
file.channels
Mix & masterProvenanceDataset QABroadcast & film post

Measures Number of audio channels.

Reading it 1 = mono, 2 = stereo.

Why A mono file makes every stereo lens meaningless, which is worth saying out loud rather than showing a flat line.

Bit depth
file.bitDepth
ProvenanceDataset QABroadcast & film post

Measures Estimated bit depth from the sample values.

Reading it 16-bit or 24-bit; an estimate, not a container reading.

Why Tells you the noise floor the format itself imposes.

Format
file.format
ProvenanceBroadcast & film post

Measures Container format as reported by the caller.

Reading it wav, mp3, m4a and so on.

Why The container says nothing about what happened before it; the bandwidth lens does.

File size
file.sizeKb
Provenance

Measures Size on disk in kilobytes.

Reading it Compare against duration to infer the bitrate.

Why Cheap sanity check: a small file with a long duration was compressed hard.

03 Loudness

How loud it is by the standards platforms actually apply.

Integrated loudness
loudness.integrated
Voice & podcastMix & masterDataset QABroadcast & film post

Measures Loudness over the whole file, BS.1770-4 gated.

Reading it Streaming targets sit near −14 LUFS, broadcast near −23, podcasts near −16.

Why The single number platforms use to turn your audio up or down, so it decides how you are heard.

Short-term max
loudness.shortTermMax
Mix & masterBroadcast & film post

Measures Loudest 3-second window.

Reading it Far above the integrated value means one passage dominates.

Why Finds the moment that will make a listener reach for the volume.

Loudness range
loudness.range
Mix & masterVoice & podcastBroadcast & film post

Measures Spread between the quiet and loud parts, EBU 3342.

Reading it Under 3 LU is very even; over 15 LU is a wide dynamic piece.

Why Predicts whether the quiet parts survive in a car or on a phone.

True peak
loudness.truePeak
Mix & masterDataset QABroadcast & film post

Measures Peak of the reconstructed waveform, 4× oversampled.

Reading it Above −1 dBTP a lossy encoder can clip what looked clean.

Why Sample peak can read fine while the played-back signal distorts; this is the number that catches it.

04 Levels

Peaks, averages, and the distance between them.

Peak
level.peak
Mix & masterDataset QABroadcast & film post

Measures Highest sample value.

Reading it 0 dBFS is the ceiling; at exactly 0 the file is probably already damaged.

Why The crudest headroom check there is, and still the first one worth doing.

RMS level
level.rms
Mix & masterDataset QA

Measures Average energy over the file.

Reading it Closer to perceived level than peak, but blind to frequency weighting.

Why Baseline for the crest factor, and the fastest read on "is this too quiet".

Crest factor
level.crest
Mix & masterDataset QABroadcast & film post

Measures Distance between peak and average.

Reading it Above 14 dB it breathes; below 8 dB a compressor or limiter flattened it.

Why Tells you whether the life was mastered out of it, independent of how loud it is.

05 Stereo

What survives when someone listens on one speaker.

Stereo correlation
stereo.correlation
Mix & masterBroadcast & film post

Measures How alike the left and right channels are, −1 to +1.

Reading it +1 is mono, around +0.5 is a normal wide mix, below 0 the channels fight each other.

Why Negative correlation hollows out or disappears when someone listens in mono, which most phones and speakers do.

Stereo width
stereo.width
Mix & master

Measures How much energy sits in the side signal rather than the middle.

Reading it 0 is fully mono; above about 0.4 is a deliberately wide image.

Why Says how much of the mix you lose the moment it is folded to mono.

Mono compatible
stereo.monoCompatible
Mix & masterBroadcast & film post

Measures Whether the mix survives a fold to mono.

Reading it False means audible cancellation.

Why One boolean that saves a listener wondering why the vocal vanished on a kitchen speaker.

06 Spectrum

Where the energy sits across frequency.

Spectrum
spectrum.curve
Mix & masterProvenanceEnvironment

Measures Average energy per frequency over the file.

Reading it Read the slope: a hump is a resonance, a cliff at the top is a codec.

Why The tonal fingerprint of the recording in one picture.

Spectral centroid
spectrum.centroid
Mix & masterEnvironment

Measures Where the weight of the spectrum sits.

Reading it Speech lands roughly 500–2000 Hz; cymbals and sibilance push it far higher.

Why One number for "bright or dark", comparable between takes and microphones.

Spectral rolloff
spectrum.rolloff
Mix & master

Measures Frequency below which 85% of the energy sits.

Reading it Moves with brightness but is less sensitive to a single loud high frequency.

Why A steadier brightness measure than the centroid when there is percussion in the file.

Spectral flatness
spectrum.flatness
EnvironmentProvenance

Measures How noise-like the spectrum is, 0 to 1.

Reading it Near 0 is a clear pitch; near 1 is broadband noise.

Why Separates a tone from a hiss without needing to know what either one is.

Tonal balance
spectrum.bands
Mix & master

Measures Energy in seven named bands from sub to air.

Reading it Compare the bars against each other, not against an absolute.

Why Turns "it sounds thin" into which band is actually missing.

07 Issues

The two complaints listeners make without knowing why.

Harshness
issue.harshness
Mix & masterVoice & podcast

Measures Excess energy in the 2–5 kHz range where the ear is most sensitive.

Reading it High means fatiguing after a few minutes.

Why Catches the thing listeners describe as tiring but cannot point at.

Muddiness
issue.muddiness
Mix & masterVoice & podcast

Measures Build-up around 200–500 Hz.

Reading it High means the mix lacks definition.

Why The most common self-produced recording problem, and the easiest to fix once named.

08 Voice

Speech, pacing, levels, and what is behind them.

Speech ratio
voice.speechRatio
Voice & podcastDataset QABroadcast & film post

Measures Share of the file where someone is talking.

Reading it Near 1 means wall-to-wall speech; near 0 means mostly silence or noise.

Why Sets what the file is for. A 0.2 ratio in a dataset take is usually a failed recording.

Speech duration
voice.speechSec
Voice & podcastDataset QA

Measures Total voiced time.

Reading it Seconds of actual talking.

Why What you are paying a transcriber or a recogniser for.

Silence duration
voice.silenceSec
Voice & podcastDataset QA

Measures Total unvoiced time.

Reading it Seconds between and around the phrases.

Why Where every background measurement has to be taken, so it bounds their reliability.

Speech segments
voice.segments
Voice & podcastDataset QA

Measures Start and end of every detected phrase.

Reading it One row per phrase.

Why The unit everything per-phrase hangs off, and the thing an editor actually clicks.

Syllable rate
voice.syllableRate
Voice & podcast

Measures Roughly how many separate bursts of speech per second, counted from dips in the envelope.

Reading it Only gaps deep enough to read as silence are counted, so this tracks word and phrase pace rather than every syllable. Across twelve reference reads it lands between 2 and 5.

Why Pace is the first thing listeners complain about, and the speaker never hears it.

Speaking rate
voice.speakingRate
Voice & podcast

Measures The syllable rate as a word.

Reading it Slow, Normal, Fast or Very fast.

Why A label travels further than a number in a report someone else reads.

Long pauses
voice.longSilences
Voice & podcast

Measures Number of gaps longer than 1.5 seconds.

Reading it A count, not a judgement.

Why Points an editor straight at the cuts worth making.

Voice peak
voice.peak
Voice & podcastDataset QA

Measures Loudest sample in the recording.

Reading it Above −1 dBFS the take is at risk.

Why The recording-level check, before anything else can be trusted.

Speech level
voice.speechLevel
Voice & podcastDataset QABroadcast & film post

Measures Average level during speech only.

Reading it A comfortable take sits near −20 dBFS; below −35 the mic was too far or too quiet.

Why Unlike the file average, this is not dragged down by the silence between phrases.

Noise floor
voice.noiseFloor
Voice & podcastDataset QAEnvironmentBroadcast & film post

Measures Level of the quietest parts.

Reading it Lower is quieter. A treated room sits far below a normal one, and a step upward means something switched on.

Why Everything you cannot remove later lives here.

Signal to noise
voice.snr
Voice & podcastDataset QABroadcast & film post

Measures Speech level against the noise floor, measured globally.

Reading it Above 30 dB is clean; below 15 dB a recogniser starts losing words.

Why The headline number, though the segmental version below is the honest one.

Segmental SNR
voice.segmentalSnr
Voice & podcastDataset QA

Measures Every phrase measured against the noise immediately around it, then taken as a median.

Reading it Same scale as the global SNR, but it does not flatter a take with a changing background.

Why One global average hides exactly what matters: a TV that comes on halfway, a passing car, someone talking next door.

SNR spread
voice.segmentalSnrSpread
Voice & podcastDataset QA

Measures Difference between the best and worst phrases (p90 minus p10).

Reading it Small means a consistent background; large means it changed during the take.

Why Decides whether the fix is one filter for the file or a hunt for individual phrases.

SNR per phrase
voice.segmentSnr
Voice & podcastDataset QA

Measures The segmental SNR for each phrase separately.

Reading it Aligned with the segment list; null where the phrase was too short to measure.

Why Turns a file-level complaint into a timestamp you can jump to.

Clipped samples
voice.clipping
Voice & podcastDataset QA

Measures Samples at or above digital full scale.

Reading it Any count above a handful is audible damage.

Why The one defect no processing can undo. It means re-record, not repair.

Gap level
voice.gapLevel
Voice & podcastDataset QAEnvironment

Measures How loud the "silence" between phrases really is.

Reading it Read it against the speech level. The smaller the distance, the more the background competes; a comfortable take leaves a wide gap.

Why The room is only quiet if you measured it, and this measures it where the speaker is not talking.

Background speech rhythm
voice.gapModulation
Voice & podcastDataset QAEnvironment

Measures Share of energy in the gaps that moves at 2–8 Hz, the rate of syllables.

Reading it High means something in the background has the cadence of talking.

Why A television or a second person in the room reads as ordinary noise on every level meter, but not on this one.

Background clatter
voice.gapModulationHigh
Voice & podcastEnvironment

Measures Share of gap energy moving faster than 8 Hz.

Reading it High means impulsive noise: keyboards, cutlery, footsteps.

Why Guards the lens above. Clatter is fast and speech is not, so the pair tells them apart.

Background cadence
voice.gapPeriodicity
Voice & podcastEnvironment

Measures How regularly the gap energy repeats at syllable spacing.

Reading it High means a steady rhythm rather than random bursts. Weak on its own: a speaker’s own breath tail is more periodic than a distant television, and broadcast compression makes radio almost aperiodic.

Why Kept as a supporting number rather than a test. Measured across a corpus it does not separate background speech by itself; it earns its place only in combination with the levels around it.

Room reverberation
voice.roomEcho
Voice & podcastDataset QA

Measures How much energy lingers after each phrase ends.

Reading it Low is a treated room; high is a kitchen or a stairwell.

Why Reverb cannot be removed afterwards in any honest way, so it has to be caught while the microphone is still up.

Room character
voice.roomEchoLabel
Voice & podcast

Measures The reverberation score as a word.

Reading it Dry, Tight, Live or Reverberant.

Why What you write in a report when the reader does not want a number.

Voice brightness
voice.brightness
Voice & podcast

Measures Spectral centre of the speech, normalised against a typical voice.

Reading it Low is muffled or off-axis; high is thin or too close to a bright microphone.

Why Catches a badly placed microphone that every level meter calls perfect.

Sibilance risk
voice.sibilance
Voice & podcast

Measures Energy concentration in the 5–9 kHz range.

Reading it Low, Medium or High.

Why Harsh S sounds survive every later processing step and make long listening painful.

Cut off at the start
voice.startsInSpeech
Voice & podcastDataset QABroadcast & film post

Measures Whether speech is already running on the very first frame.

Reading it True means the recording started too late.

Why A missing first word is invisible in every waveform view and fatal in a dataset.

Cut off at the end
voice.endsInSpeech
Voice & podcastDataset QABroadcast & film post

Measures Whether speech is still running on the last frame.

Reading it True means the recording stopped too early.

Why Same as above, at the other end, and just as easy to miss.

Level at the start
voice.edgeStartDb
Dataset QABroadcast & film post

Measures Envelope level on the first frames.

Reading it Loud means a hard cut; quiet means a natural start.

Why Separates a truncated take from one that simply begins softly.

Level at the end
voice.edgeEndDb
Dataset QABroadcast & film post

Measures Envelope level on the last frames.

Reading it Loud means a hard cut; quiet means it faded out.

Why The counterpart at the end, same reasoning.

09 Per phrase

The same questions, asked of one sentence at a time.

Phrase level
segment.level
Voice & podcastDataset QA

Measures Speech level of this phrase alone.

Reading it Compare phrases against each other to spot a speaker drifting off the microphone.

Why A take can average perfectly while half the phrases are too quiet.

Phrase clipping
segment.clipFraction
Voice & podcastDataset QA

Measures Share of clipped samples inside this phrase.

Reading it Anything above zero points at the phrase to re-record.

Why Narrows "this file clips" down to which sentence.

Phrase reverb
segment.decayScore
Voice & podcastDataset QA

Measures How much energy lingers after this phrase, relative to the phrase itself.

Reading it Null when nothing usable follows, which is honest rather than zero.

Why Reverb varies within a take when a speaker turns their head; the file average hides that.

Decay slope
segment.tailDecay
Voice & podcastDataset QA

Measures How fast the sound falls away after this phrase, in dB per second.

Reading it Steeper is drier, shallower rings. Needs a quiet gap behind the phrase, which across a corpus only about one phrase in ten has; read it as a sample, not as an average.

Why A physical measure of the room rather than a score, so it can be compared between recordings and devices.

Background character
segment.flankFlatness
Voice & podcastDataset QAEnvironment

Measures Spectral flatness of the pauses around this phrase.

Reading it High is broadband room noise; low means something tonal is playing underneath.

Why Two phrases with the same poor SNR need different fixes, and this says which is which.

10 Content

What kind of audio this is, and how firmly we hold that.

Content type
signal.contentType
Voice & podcastMix & masterProvenanceDataset QAEnvironment

Measures Whether this is voice, music, mixed, noise or silence.

Reading it "Unknown" is a real answer: the evidence did not agree.

Why Decides which of the other lenses mean anything at all for this file.

Content confidence
signal.contentConfidence
Dataset QAProvenance

Measures How firmly the content type is held.

Reading it Capped at 0.5 for files under three seconds.

Why A verdict without its confidence is a guess dressed as a fact.

Evidence
signal.contentEvidence
Dataset QAProvenance

Measures Which independent checks back the content type up.

Reading it A list of short facts, such as a beat grid or a centred stereo image.

Why Makes the verdict arguable instead of oracular, which is the only way anyone trusts it twice.

Beat stability
signal.beatStability
Mix & masterDataset QA

Measures Share of 6-second windows agreeing on one tempo.

Reading it Null below 24 seconds, because a claim needs enough windows to be worth making.

Why Music holds a grid and speech does not, which is the most reliable single separator we found.

Spectral flatness (opening)
signal.tonality
EnvironmentDataset QA

Measures How noise-like the opening of the file is, 0 to 1.

Reading it Near zero is a pure tone: an alarm, a test signal, a beep. Near one is broadband noise.

Why Speech is never a pure tone, so a near-zero reading rules it out on its own.

Brightness
signal.brightnessBucket
Mix & master

Measures The overall tone as a word.

Reading it Dark, Balanced, Bright or Very bright.

Why Sorting a library by feel needs words, not hertz.

Dynamics
signal.dynamicsBucket
Mix & master

Measures The crest factor as a word.

Reading it Flat, Compressed, Modern or Dynamic.

Why Tells you at a glance whether a file has been mastered already.

Dominant band
signal.dominantBand
Mix & masterEnvironment

Measures Which frequency band carries the most energy.

Reading it Named, from sub to air.

Why The shortest possible description of what a file sounds like.

Monophonic probability
signal.monophonic
ProvenanceDataset QA

Measures How likely it is that both channels carry the same thing.

Reading it Near 1 means the stereo file is mono in disguise.

Why A stereo container proves nothing; this catches an upmixed or duplicated channel.

DC offset
signal.dcOffset
Dataset QAProvenanceBroadcast & film post

Measures Constant shift of the waveform away from zero.

Reading it Anything meaningfully above zero points at faulty hardware.

Why Costs headroom and shows up nowhere else. One number for the whole file is enough.

Has clipping
signal.clipping
Dataset QAMix & masterBroadcast & film post

Measures Whether the file clips anywhere.

Reading it A yes or no for triage.

Why The first filter when sorting a folder of hundreds.

Clipping regions
signal.clippingRegions
Dataset QAMix & masterBroadcast & film post

Measures How many separate stretches clip.

Reading it One region is an accident; twenty is a recording-level problem.

Why Distinguishes a single loud moment from a take recorded far too hot.

Silent regions
signal.silenceRegions
Dataset QABroadcast & film post

Measures How many separate silent stretches there are.

Reading it A count.

Why Many short silences is a stuttering edit; one long one is dead air.

Total silence
signal.silenceTotalSec
Dataset QABroadcast & film post

Measures How much of the file is silent.

Reading it Seconds.

Why Tells you how much of what you are storing and paying for is nothing.

Timeline regions
signal.regions
Voice & podcastProvenanceDataset QAEnvironment

Measures The file cut into stretches labelled voice, silence, music, noise or clipping.

Reading it Each region carries its own confidence.

Why The map you navigate by, and the thing an agent can act on region by region.

Tags
signal.tags
Dataset QA

Measures Short keywords summarising the file.

Reading it Meant for search, not for reading.

Why Makes a library findable without opening anything.

11 Over time

Values that move, returned as curves.

Loudness over time
series.shortTermLufs
Mix & masterVoice & podcastBroadcast & film post

Measures Short-term loudness in 400 ms steps.

Reading it Flat = evenly loud; jagged = the level moves.

Why Shows where the loudness range comes from instead of only how big it is.

Waveform
series.waveform
Voice & podcastMix & masterProvenanceDataset QAEnvironmentBroadcast & film post

Measures Downsampled envelope of the signal.

Reading it The familiar shape; loud is tall.

Why Orientation. Everything else is easier to read once you know where you are.

Voice envelope
series.rmsEnvelope
Voice & podcast

Measures Level over time at about 50 frames per second.

Reading it The shape of the speaking, finer than the waveform view.

Why Fine enough to see individual syllables, which is where pacing and clipping live.

Voice activity mask
series.voicedMask
Voice & podcastDataset QA

Measures Speech or not, for every envelope frame.

Reading it Aligned one-to-one with the envelope.

Why Lay it over any other lens and every measurement splits into "during speech" and "between speech".

Energy envelope
series.energyEnvelope
ProvenanceDataset QAEnvironment

Measures Level over time, fixed resolution regardless of file length.

Reading it The overview shape.

Why Comparable between files of different lengths, which the waveform is not.

Brightness envelope
series.centroidEnvelope
ProvenanceEnvironment

Measures Spectral centre over time, same resolution as the energy envelope.

Reading it Rising means the material got brighter.

Why Brightness moves where level does not: an EQ change, a different microphone, another source entirely.

12 Lanes

Block-by-block measurements you can stack on one time axis.

Spectrum lane
grid.terrain
Voice & podcastMix & masterProvenanceDataset QAEnvironment

Measures Energy per frequency band over time, 48 log-spaced bands.

Reading it Bright is loud. Vertical stripes are hits; horizontal bars are sustained tones.

Why The shape of the sound at a glance, and the layer every other lane is read against.

Level lane
grid.level
Voice & podcastMix & masterDataset QABroadcast & film post

Measures RMS level per block.

Reading it Higher is louder; a flat line is compressed or steady material.

Why Finds the loud and quiet stretches without scrubbing through the file.

Peak lane
grid.peak
Mix & masterDataset QA

Measures Highest sample per block.

Reading it Touching 0 dBFS marks the blocks that clip.

Why Pairs with the level lane to give the crest factor its two halves.

Dynamics lane
grid.crest
Mix & masterProvenance

Measures Crest factor per block.

Reading it Above 14 dB it breathes; below 8 dB it was squashed.

Why A flat stretch inside an otherwise dynamic file is almost always a different source.

Brightness lane
grid.centroid
Mix & masterProvenanceEnvironment

Measures Spectral centre per block.

Reading it Rising is brighter, falling is darker or duller.

Why Follows tonal changes the spectrum lane shows but does not summarise.

Bandwidth lane
grid.bandwidth
ProvenanceDataset QA

Measures Where the spectrum closes off, block by block. Null when too quiet to measure.

Reading it Measured: 64 kbps sits near 11 kHz, 128 near 16.5, 256 near 19.

Why Reads the file’s history: how hard it was compressed, and exactly where two sources were spliced together.

Noise floor lane
grid.noiseFloor
Voice & podcastDataset QAEnvironment

Measures The quietest frames near each block, as a 25th percentile over a ±3 block window.

Reading it A step up means noise arrived. Digital silence sits near −120 dB.

Why Noise under speech costs a recogniser words, and this shows the second it starts.

Stereo lane
grid.correlation
Mix & master

Measures Left against right per block, on a fixed −1 to +1 scale.

Reading it +1 is mono, lower is wider, below 0 the channels fight.

Why Catches a phase problem in one passage that the file-level average washes out.

Width lane
grid.width
Mix & master

Measures Stereo width per block.

Reading it Drawn as the area behind the correlation line.

Why Shows where a mix opens up and where it collapses to the centre.

Flatness lane
grid.flatness
EnvironmentProvenance

Measures How noise-like each block is, 0 to 1.

Reading it Low is a clear pitch; high is broadband.

Why Separates a beep from a voice and hum from room noise, over time.

Clipping lane
grid.clipFraction
Mix & masterDataset QABroadcast & film post

Measures Share of samples at or over full scale per block.

Reading it A bar means clipping happened there.

Why Damage you cannot undo, located to the second.

DC lane
grid.dcOffset
Provenance

Measures Waveform shift away from zero per block.

Reading it Should sit flat at zero for the whole file.

Why A step means the hardware changed mid-recording, which nothing else here would reveal.

13 Planned

Combinations of the above, written down before they are built.

Headroom lane
derived.headroom
Voice & podcastDataset QAEnvironmentBroadcast & film post

Measures Distance between the content and the local noise floor, block by block.

Reading it Above 30 dB is clean; below 15 dB a recogniser starts losing words.

Why This is the number that decides whether audio is usable, and right now it is spread over two lanes the reader has to subtract by eye.

Novelty lane
derived.novelty
ProvenanceDataset QA

Measures How much the material changes across each moment, comparing the two seconds before with the two after over bandwidth, noise floor, brightness, crest and stereo correlation together.

Reading it Peaks mark boundaries. Read all of them, not just the highest; on dense music the tallest peak is not always the one you were looking for.

Why Points at where something changes without being told what to look for. It carries stereo correlation, so it is the one that catches a channel collapse; the change lane reads the mono terrain and catches a splice instead. Between them the five test cases are covered. Scaled so it never saturates: the top of the lane stays available for the genuinely exceptional.

Change lane
derived.flux
Mix & masterProvenance

Measures How much the spectrum changes from one block to the next.

Reading it Peaks are onsets, edits and transitions.

Why The terrain is already there and never differentiated, and its derivative is where the edit points live. It found the splice at exactly 15.0 s on the test file that novelty missed. Blind to a stereo change by construction, since the terrain is mono.

Noise character lane
derived.bandFlatness
EnvironmentVoice & podcast

Measures Tonality measured separately below 250 Hz, from 250 Hz to 4 kHz, and above.

Reading it Tonal low is hum, flat high is hiss, tonal mid is music or a television. Three lines, one per band group.

Why One flatness number per block can say how much noise there is but never what kind, and the fix depends entirely on the kind.