/lenses
Everything we can see.
93 measurements, each one stating what it measures, how to read it, and why it is worth measuring. A lens that cannot answer the third question does not ship.
Nobody needs all 93 at once. Each profile is a working set: choose one to see only the lenses that answer its question.
What the container says, and what it cannot tell you.
- Duration file.durationVoice & podcastMix & masterProvenanceDataset QAEnvironmentBroadcast & film post
-
Measures Length of the audio.
Reading it Seconds.
Why Sets what else can be claimed: below three seconds most verdicts are guesses.
- Sample rate file.sampleRateProvenanceDataset QABroadcast & film post
-
Measures Samples per second.
Reading it 44.1k and 48k are standard; 8k or 16k means it came through a phone or a codec.
Why Caps the highest frequency that can exist at all, so it bounds every spectral reading.
- Channels file.channelsMix & masterProvenanceDataset QABroadcast & film post
-
Measures Number of audio channels.
Reading it 1 = mono, 2 = stereo.
Why A mono file makes every stereo lens meaningless, which is worth saying out loud rather than showing a flat line.
- Bit depth file.bitDepthProvenanceDataset QABroadcast & film post
-
Measures Estimated bit depth from the sample values.
Reading it 16-bit or 24-bit; an estimate, not a container reading.
Why Tells you the noise floor the format itself imposes.
- Format file.formatProvenanceBroadcast & film post
-
Measures Container format as reported by the caller.
Reading it wav, mp3, m4a and so on.
Why The container says nothing about what happened before it; the bandwidth lens does.
- File size file.sizeKbProvenance
-
Measures Size on disk in kilobytes.
Reading it Compare against duration to infer the bitrate.
Why Cheap sanity check: a small file with a long duration was compressed hard.
How loud it is by the standards platforms actually apply.
- Integrated loudness loudness.integratedVoice & podcastMix & masterDataset QABroadcast & film post
-
Measures Loudness over the whole file, BS.1770-4 gated.
Reading it Streaming targets sit near −14 LUFS, broadcast near −23, podcasts near −16.
Why The single number platforms use to turn your audio up or down, so it decides how you are heard.
- Short-term max loudness.shortTermMaxMix & masterBroadcast & film post
-
Measures Loudest 3-second window.
Reading it Far above the integrated value means one passage dominates.
Why Finds the moment that will make a listener reach for the volume.
- Loudness range loudness.rangeMix & masterVoice & podcastBroadcast & film post
-
Measures Spread between the quiet and loud parts, EBU 3342.
Reading it Under 3 LU is very even; over 15 LU is a wide dynamic piece.
Why Predicts whether the quiet parts survive in a car or on a phone.
- True peak loudness.truePeakMix & masterDataset QABroadcast & film post
-
Measures Peak of the reconstructed waveform, 4× oversampled.
Reading it Above −1 dBTP a lossy encoder can clip what looked clean.
Why Sample peak can read fine while the played-back signal distorts; this is the number that catches it.
Peaks, averages, and the distance between them.
- Peak level.peakMix & masterDataset QABroadcast & film post
-
Measures Highest sample value.
Reading it 0 dBFS is the ceiling; at exactly 0 the file is probably already damaged.
Why The crudest headroom check there is, and still the first one worth doing.
- RMS level level.rmsMix & masterDataset QA
-
Measures Average energy over the file.
Reading it Closer to perceived level than peak, but blind to frequency weighting.
Why Baseline for the crest factor, and the fastest read on "is this too quiet".
- Crest factor level.crestMix & masterDataset QABroadcast & film post
-
Measures Distance between peak and average.
Reading it Above 14 dB it breathes; below 8 dB a compressor or limiter flattened it.
Why Tells you whether the life was mastered out of it, independent of how loud it is.
What survives when someone listens on one speaker.
- Stereo correlation stereo.correlationMix & masterBroadcast & film post
-
Measures How alike the left and right channels are, −1 to +1.
Reading it +1 is mono, around +0.5 is a normal wide mix, below 0 the channels fight each other.
Why Negative correlation hollows out or disappears when someone listens in mono, which most phones and speakers do.
- Stereo width stereo.widthMix & master
-
Measures How much energy sits in the side signal rather than the middle.
Reading it 0 is fully mono; above about 0.4 is a deliberately wide image.
Why Says how much of the mix you lose the moment it is folded to mono.
- Mono compatible stereo.monoCompatibleMix & masterBroadcast & film post
-
Measures Whether the mix survives a fold to mono.
Reading it False means audible cancellation.
Why One boolean that saves a listener wondering why the vocal vanished on a kitchen speaker.
Where the energy sits across frequency.
- Spectrum spectrum.curveMix & masterProvenanceEnvironment
-
Measures Average energy per frequency over the file.
Reading it Read the slope: a hump is a resonance, a cliff at the top is a codec.
Why The tonal fingerprint of the recording in one picture.
- Spectral centroid spectrum.centroidMix & masterEnvironment
-
Measures Where the weight of the spectrum sits.
Reading it Speech lands roughly 500–2000 Hz; cymbals and sibilance push it far higher.
Why One number for "bright or dark", comparable between takes and microphones.
- Spectral rolloff spectrum.rolloffMix & master
-
Measures Frequency below which 85% of the energy sits.
Reading it Moves with brightness but is less sensitive to a single loud high frequency.
Why A steadier brightness measure than the centroid when there is percussion in the file.
- Spectral flatness spectrum.flatnessEnvironmentProvenance
-
Measures How noise-like the spectrum is, 0 to 1.
Reading it Near 0 is a clear pitch; near 1 is broadband noise.
Why Separates a tone from a hiss without needing to know what either one is.
- Tonal balance spectrum.bandsMix & master
-
Measures Energy in seven named bands from sub to air.
Reading it Compare the bars against each other, not against an absolute.
Why Turns "it sounds thin" into which band is actually missing.
The two complaints listeners make without knowing why.
- Harshness issue.harshnessMix & masterVoice & podcast
-
Measures Excess energy in the 2–5 kHz range where the ear is most sensitive.
Reading it High means fatiguing after a few minutes.
Why Catches the thing listeners describe as tiring but cannot point at.
- Muddiness issue.muddinessMix & masterVoice & podcast
-
Measures Build-up around 200–500 Hz.
Reading it High means the mix lacks definition.
Why The most common self-produced recording problem, and the easiest to fix once named.
Speech, pacing, levels, and what is behind them.
- Speech ratio voice.speechRatioVoice & podcastDataset QABroadcast & film post
-
Measures Share of the file where someone is talking.
Reading it Near 1 means wall-to-wall speech; near 0 means mostly silence or noise.
Why Sets what the file is for. A 0.2 ratio in a dataset take is usually a failed recording.
- Speech duration voice.speechSecVoice & podcastDataset QA
-
Measures Total voiced time.
Reading it Seconds of actual talking.
Why What you are paying a transcriber or a recogniser for.
- Silence duration voice.silenceSecVoice & podcastDataset QA
-
Measures Total unvoiced time.
Reading it Seconds between and around the phrases.
Why Where every background measurement has to be taken, so it bounds their reliability.
- Speech segments voice.segmentsVoice & podcastDataset QA
-
Measures Start and end of every detected phrase.
Reading it One row per phrase.
Why The unit everything per-phrase hangs off, and the thing an editor actually clicks.
- Syllable rate voice.syllableRateVoice & podcast
-
Measures Roughly how many separate bursts of speech per second, counted from dips in the envelope.
Reading it Only gaps deep enough to read as silence are counted, so this tracks word and phrase pace rather than every syllable. Across twelve reference reads it lands between 2 and 5.
Why Pace is the first thing listeners complain about, and the speaker never hears it.
- Speaking rate voice.speakingRateVoice & podcast
-
Measures The syllable rate as a word.
Reading it Slow, Normal, Fast or Very fast.
Why A label travels further than a number in a report someone else reads.
- Long pauses voice.longSilencesVoice & podcast
-
Measures Number of gaps longer than 1.5 seconds.
Reading it A count, not a judgement.
Why Points an editor straight at the cuts worth making.
- Voice peak voice.peakVoice & podcastDataset QA
-
Measures Loudest sample in the recording.
Reading it Above −1 dBFS the take is at risk.
Why The recording-level check, before anything else can be trusted.
- Speech level voice.speechLevelVoice & podcastDataset QABroadcast & film post
-
Measures Average level during speech only.
Reading it A comfortable take sits near −20 dBFS; below −35 the mic was too far or too quiet.
Why Unlike the file average, this is not dragged down by the silence between phrases.
- Noise floor voice.noiseFloorVoice & podcastDataset QAEnvironmentBroadcast & film post
-
Measures Level of the quietest parts.
Reading it Lower is quieter. A treated room sits far below a normal one, and a step upward means something switched on.
Why Everything you cannot remove later lives here.
- Signal to noise voice.snrVoice & podcastDataset QABroadcast & film post
-
Measures Speech level against the noise floor, measured globally.
Reading it Above 30 dB is clean; below 15 dB a recogniser starts losing words.
Why The headline number, though the segmental version below is the honest one.
- Segmental SNR voice.segmentalSnrVoice & podcastDataset QA
-
Measures Every phrase measured against the noise immediately around it, then taken as a median.
Reading it Same scale as the global SNR, but it does not flatter a take with a changing background.
Why One global average hides exactly what matters: a TV that comes on halfway, a passing car, someone talking next door.
- SNR spread voice.segmentalSnrSpreadVoice & podcastDataset QA
-
Measures Difference between the best and worst phrases (p90 minus p10).
Reading it Small means a consistent background; large means it changed during the take.
Why Decides whether the fix is one filter for the file or a hunt for individual phrases.
- SNR per phrase voice.segmentSnrVoice & podcastDataset QA
-
Measures The segmental SNR for each phrase separately.
Reading it Aligned with the segment list; null where the phrase was too short to measure.
Why Turns a file-level complaint into a timestamp you can jump to.
- Clipped samples voice.clippingVoice & podcastDataset QA
-
Measures Samples at or above digital full scale.
Reading it Any count above a handful is audible damage.
Why The one defect no processing can undo. It means re-record, not repair.
- Gap level voice.gapLevelVoice & podcastDataset QAEnvironment
-
Measures How loud the "silence" between phrases really is.
Reading it Read it against the speech level. The smaller the distance, the more the background competes; a comfortable take leaves a wide gap.
Why The room is only quiet if you measured it, and this measures it where the speaker is not talking.
- Background speech rhythm voice.gapModulationVoice & podcastDataset QAEnvironment
-
Measures Share of energy in the gaps that moves at 2–8 Hz, the rate of syllables.
Reading it High means something in the background has the cadence of talking.
Why A television or a second person in the room reads as ordinary noise on every level meter, but not on this one.
- Background clatter voice.gapModulationHighVoice & podcastEnvironment
-
Measures Share of gap energy moving faster than 8 Hz.
Reading it High means impulsive noise: keyboards, cutlery, footsteps.
Why Guards the lens above. Clatter is fast and speech is not, so the pair tells them apart.
- Background cadence voice.gapPeriodicityVoice & podcastEnvironment
-
Measures How regularly the gap energy repeats at syllable spacing.
Reading it High means a steady rhythm rather than random bursts. Weak on its own: a speaker’s own breath tail is more periodic than a distant television, and broadcast compression makes radio almost aperiodic.
Why Kept as a supporting number rather than a test. Measured across a corpus it does not separate background speech by itself; it earns its place only in combination with the levels around it.
- Room reverberation voice.roomEchoVoice & podcastDataset QA
-
Measures How much energy lingers after each phrase ends.
Reading it Low is a treated room; high is a kitchen or a stairwell.
Why Reverb cannot be removed afterwards in any honest way, so it has to be caught while the microphone is still up.
- Room character voice.roomEchoLabelVoice & podcast
-
Measures The reverberation score as a word.
Reading it Dry, Tight, Live or Reverberant.
Why What you write in a report when the reader does not want a number.
- Voice brightness voice.brightnessVoice & podcast
-
Measures Spectral centre of the speech, normalised against a typical voice.
Reading it Low is muffled or off-axis; high is thin or too close to a bright microphone.
Why Catches a badly placed microphone that every level meter calls perfect.
- Sibilance risk voice.sibilanceVoice & podcast
-
Measures Energy concentration in the 5–9 kHz range.
Reading it Low, Medium or High.
Why Harsh S sounds survive every later processing step and make long listening painful.
- Cut off at the start voice.startsInSpeechVoice & podcastDataset QABroadcast & film post
-
Measures Whether speech is already running on the very first frame.
Reading it True means the recording started too late.
Why A missing first word is invisible in every waveform view and fatal in a dataset.
- Cut off at the end voice.endsInSpeechVoice & podcastDataset QABroadcast & film post
-
Measures Whether speech is still running on the last frame.
Reading it True means the recording stopped too early.
Why Same as above, at the other end, and just as easy to miss.
- Level at the start voice.edgeStartDbDataset QABroadcast & film post
-
Measures Envelope level on the first frames.
Reading it Loud means a hard cut; quiet means a natural start.
Why Separates a truncated take from one that simply begins softly.
- Level at the end voice.edgeEndDbDataset QABroadcast & film post
-
Measures Envelope level on the last frames.
Reading it Loud means a hard cut; quiet means it faded out.
Why The counterpart at the end, same reasoning.
The same questions, asked of one sentence at a time.
- Phrase level segment.levelVoice & podcastDataset QA
-
Measures Speech level of this phrase alone.
Reading it Compare phrases against each other to spot a speaker drifting off the microphone.
Why A take can average perfectly while half the phrases are too quiet.
- Phrase clipping segment.clipFractionVoice & podcastDataset QA
-
Measures Share of clipped samples inside this phrase.
Reading it Anything above zero points at the phrase to re-record.
Why Narrows "this file clips" down to which sentence.
- Phrase reverb segment.decayScoreVoice & podcastDataset QA
-
Measures How much energy lingers after this phrase, relative to the phrase itself.
Reading it Null when nothing usable follows, which is honest rather than zero.
Why Reverb varies within a take when a speaker turns their head; the file average hides that.
- Decay slope segment.tailDecayVoice & podcastDataset QA
-
Measures How fast the sound falls away after this phrase, in dB per second.
Reading it Steeper is drier, shallower rings. Needs a quiet gap behind the phrase, which across a corpus only about one phrase in ten has; read it as a sample, not as an average.
Why A physical measure of the room rather than a score, so it can be compared between recordings and devices.
- Background character segment.flankFlatnessVoice & podcastDataset QAEnvironment
-
Measures Spectral flatness of the pauses around this phrase.
Reading it High is broadband room noise; low means something tonal is playing underneath.
Why Two phrases with the same poor SNR need different fixes, and this says which is which.
What kind of audio this is, and how firmly we hold that.
- Content type signal.contentTypeVoice & podcastMix & masterProvenanceDataset QAEnvironment
-
Measures Whether this is voice, music, mixed, noise or silence.
Reading it "Unknown" is a real answer: the evidence did not agree.
Why Decides which of the other lenses mean anything at all for this file.
- Content confidence signal.contentConfidenceDataset QAProvenance
-
Measures How firmly the content type is held.
Reading it Capped at 0.5 for files under three seconds.
Why A verdict without its confidence is a guess dressed as a fact.
- Evidence signal.contentEvidenceDataset QAProvenance
-
Measures Which independent checks back the content type up.
Reading it A list of short facts, such as a beat grid or a centred stereo image.
Why Makes the verdict arguable instead of oracular, which is the only way anyone trusts it twice.
- Beat stability signal.beatStabilityMix & masterDataset QA
-
Measures Share of 6-second windows agreeing on one tempo.
Reading it Null below 24 seconds, because a claim needs enough windows to be worth making.
Why Music holds a grid and speech does not, which is the most reliable single separator we found.
- Spectral flatness (opening) signal.tonalityEnvironmentDataset QA
-
Measures How noise-like the opening of the file is, 0 to 1.
Reading it Near zero is a pure tone: an alarm, a test signal, a beep. Near one is broadband noise.
Why Speech is never a pure tone, so a near-zero reading rules it out on its own.
- Brightness signal.brightnessBucketMix & master
-
Measures The overall tone as a word.
Reading it Dark, Balanced, Bright or Very bright.
Why Sorting a library by feel needs words, not hertz.
- Dynamics signal.dynamicsBucketMix & master
-
Measures The crest factor as a word.
Reading it Flat, Compressed, Modern or Dynamic.
Why Tells you at a glance whether a file has been mastered already.
- Dominant band signal.dominantBandMix & masterEnvironment
-
Measures Which frequency band carries the most energy.
Reading it Named, from sub to air.
Why The shortest possible description of what a file sounds like.
- Monophonic probability signal.monophonicProvenanceDataset QA
-
Measures How likely it is that both channels carry the same thing.
Reading it Near 1 means the stereo file is mono in disguise.
Why A stereo container proves nothing; this catches an upmixed or duplicated channel.
- DC offset signal.dcOffsetDataset QAProvenanceBroadcast & film post
-
Measures Constant shift of the waveform away from zero.
Reading it Anything meaningfully above zero points at faulty hardware.
Why Costs headroom and shows up nowhere else. One number for the whole file is enough.
- Has clipping signal.clippingDataset QAMix & masterBroadcast & film post
-
Measures Whether the file clips anywhere.
Reading it A yes or no for triage.
Why The first filter when sorting a folder of hundreds.
- Clipping regions signal.clippingRegionsDataset QAMix & masterBroadcast & film post
-
Measures How many separate stretches clip.
Reading it One region is an accident; twenty is a recording-level problem.
Why Distinguishes a single loud moment from a take recorded far too hot.
- Silent regions signal.silenceRegionsDataset QABroadcast & film post
-
Measures How many separate silent stretches there are.
Reading it A count.
Why Many short silences is a stuttering edit; one long one is dead air.
- Total silence signal.silenceTotalSecDataset QABroadcast & film post
-
Measures How much of the file is silent.
Reading it Seconds.
Why Tells you how much of what you are storing and paying for is nothing.
- Timeline regions signal.regionsVoice & podcastProvenanceDataset QAEnvironment
-
Measures The file cut into stretches labelled voice, silence, music, noise or clipping.
Reading it Each region carries its own confidence.
Why The map you navigate by, and the thing an agent can act on region by region.
- Tags signal.tagsDataset QA
-
Measures Short keywords summarising the file.
Reading it Meant for search, not for reading.
Why Makes a library findable without opening anything.
Values that move, returned as curves.
- Loudness over time series.shortTermLufsMix & masterVoice & podcastBroadcast & film post
-
Measures Short-term loudness in 400 ms steps.
Reading it Flat = evenly loud; jagged = the level moves.
Why Shows where the loudness range comes from instead of only how big it is.
- Waveform series.waveformVoice & podcastMix & masterProvenanceDataset QAEnvironmentBroadcast & film post
-
Measures Downsampled envelope of the signal.
Reading it The familiar shape; loud is tall.
Why Orientation. Everything else is easier to read once you know where you are.
- Voice envelope series.rmsEnvelopeVoice & podcast
-
Measures Level over time at about 50 frames per second.
Reading it The shape of the speaking, finer than the waveform view.
Why Fine enough to see individual syllables, which is where pacing and clipping live.
- Voice activity mask series.voicedMaskVoice & podcastDataset QA
-
Measures Speech or not, for every envelope frame.
Reading it Aligned one-to-one with the envelope.
Why Lay it over any other lens and every measurement splits into "during speech" and "between speech".
- Energy envelope series.energyEnvelopeProvenanceDataset QAEnvironment
-
Measures Level over time, fixed resolution regardless of file length.
Reading it The overview shape.
Why Comparable between files of different lengths, which the waveform is not.
- Brightness envelope series.centroidEnvelopeProvenanceEnvironment
-
Measures Spectral centre over time, same resolution as the energy envelope.
Reading it Rising means the material got brighter.
Why Brightness moves where level does not: an EQ change, a different microphone, another source entirely.
Block-by-block measurements you can stack on one time axis.
- Spectrum lane grid.terrainVoice & podcastMix & masterProvenanceDataset QAEnvironment
-
Measures Energy per frequency band over time, 48 log-spaced bands.
Reading it Bright is loud. Vertical stripes are hits; horizontal bars are sustained tones.
Why The shape of the sound at a glance, and the layer every other lane is read against.
- Level lane grid.levelVoice & podcastMix & masterDataset QABroadcast & film post
-
Measures RMS level per block.
Reading it Higher is louder; a flat line is compressed or steady material.
Why Finds the loud and quiet stretches without scrubbing through the file.
- Peak lane grid.peakMix & masterDataset QA
-
Measures Highest sample per block.
Reading it Touching 0 dBFS marks the blocks that clip.
Why Pairs with the level lane to give the crest factor its two halves.
- Dynamics lane grid.crestMix & masterProvenance
-
Measures Crest factor per block.
Reading it Above 14 dB it breathes; below 8 dB it was squashed.
Why A flat stretch inside an otherwise dynamic file is almost always a different source.
- Brightness lane grid.centroidMix & masterProvenanceEnvironment
-
Measures Spectral centre per block.
Reading it Rising is brighter, falling is darker or duller.
Why Follows tonal changes the spectrum lane shows but does not summarise.
- Bandwidth lane grid.bandwidthProvenanceDataset QA
-
Measures Where the spectrum closes off, block by block. Null when too quiet to measure.
Reading it Measured: 64 kbps sits near 11 kHz, 128 near 16.5, 256 near 19.
Why Reads the file’s history: how hard it was compressed, and exactly where two sources were spliced together.
- Noise floor lane grid.noiseFloorVoice & podcastDataset QAEnvironment
-
Measures The quietest frames near each block, as a 25th percentile over a ±3 block window.
Reading it A step up means noise arrived. Digital silence sits near −120 dB.
Why Noise under speech costs a recogniser words, and this shows the second it starts.
- Stereo lane grid.correlationMix & master
-
Measures Left against right per block, on a fixed −1 to +1 scale.
Reading it +1 is mono, lower is wider, below 0 the channels fight.
Why Catches a phase problem in one passage that the file-level average washes out.
- Width lane grid.widthMix & master
-
Measures Stereo width per block.
Reading it Drawn as the area behind the correlation line.
Why Shows where a mix opens up and where it collapses to the centre.
- Flatness lane grid.flatnessEnvironmentProvenance
-
Measures How noise-like each block is, 0 to 1.
Reading it Low is a clear pitch; high is broadband.
Why Separates a beep from a voice and hum from room noise, over time.
- Clipping lane grid.clipFractionMix & masterDataset QABroadcast & film post
-
Measures Share of samples at or over full scale per block.
Reading it A bar means clipping happened there.
Why Damage you cannot undo, located to the second.
- DC lane grid.dcOffsetProvenance
-
Measures Waveform shift away from zero per block.
Reading it Should sit flat at zero for the whole file.
Why A step means the hardware changed mid-recording, which nothing else here would reveal.
Combinations of the above, written down before they are built.
- Headroom lane derived.headroomVoice & podcastDataset QAEnvironmentBroadcast & film post
-
Measures Distance between the content and the local noise floor, block by block.
Reading it Above 30 dB is clean; below 15 dB a recogniser starts losing words.
Why This is the number that decides whether audio is usable, and right now it is spread over two lanes the reader has to subtract by eye.
- Novelty lane derived.noveltyProvenanceDataset QA
-
Measures How much the material changes across each moment, comparing the two seconds before with the two after over bandwidth, noise floor, brightness, crest and stereo correlation together.
Reading it Peaks mark boundaries. Read all of them, not just the highest; on dense music the tallest peak is not always the one you were looking for.
Why Points at where something changes without being told what to look for. It carries stereo correlation, so it is the one that catches a channel collapse; the change lane reads the mono terrain and catches a splice instead. Between them the five test cases are covered. Scaled so it never saturates: the top of the lane stays available for the genuinely exceptional.
- Change lane derived.fluxMix & masterProvenance
-
Measures How much the spectrum changes from one block to the next.
Reading it Peaks are onsets, edits and transitions.
Why The terrain is already there and never differentiated, and its derivative is where the edit points live. It found the splice at exactly 15.0 s on the test file that novelty missed. Blind to a stereo change by construction, since the terrain is mono.
- Noise character lane derived.bandFlatnessEnvironmentVoice & podcast
-
Measures Tonality measured separately below 250 Hz, from 250 Hz to 4 kHz, and above.
Reading it Tonal low is hum, flat high is hiss, tonal mid is music or a television. Three lines, one per band group.
Why One flatness number per block can say how much noise there is but never what kind, and the fix depends entirely on the kind.