JavaScript MIT

audio

High-level audio manipulations

A

audiojs

Dernière activité 27 sept. 2026
audiojs/audio

311

étoiles

10

forks

6

issues ouvertes

audioaudiojsjavascript

Ce README est souvent en anglais.

audio test npm

High-level audio manipulations: loading, playback, analysis and editing. 140+ plugins: denoise, dynamics, EQ, filters, effects, reverb, time and pitch, spatial, synthesis, music analysis.

  • Any Format — fast wasm codecs, no ffmpeg.
  • Non-destructive — virtual edits, infinite undo, instant clone.
  • Stream-first — playback/encode during decode, realtime editing.
  • Paged — no 2Gb memory limit, open 10Gb+ files.
  • Analysis — loudness, spectrum, beats, pitch, chords, key.
  • Modular – pluggable ops, tree-shakable.
  • CLI — playback, batch processing, scripting, unix pipes, tab completion.
  • Cross-platform — browsers, node, deno, bun.

npm install audio

Start   Recipes   API   CLI   FAQ   Plugins   Architecture   Comparison   RX 12

Start

Node

npm i audio

import audio from 'audio'

let song = audio('song.mp3')
song.play()                                  // the speakers, as a page plays it: no ffmpeg, no player to spawn
song.pause(); song.seek(60); song.resume()

let voice = audio('voice.wav').trim().normalize('podcast')
await voice.save('clean.mp3')
let { pass, rules } = await voice.check('podcast')  // each rule of Apple's spec, measured

Browser

<script type="module">
  import audio from 'https://esm.sh/audio'
  audio('./song.mp3').trim().normalize().fade(0.5, 2).clip({ at: 60, duration: 30 }).play()
</script>

Audio anywhere

import 'audio/polyfill'          // the page's HTMLAudioElement, in Node, Bun and Deno
let song = new Audio('song.mp3')
song.onended = () => console.log('done')
await song.play()
song.currentTime = 30; song.volume = 0.5; song.loop = true

Events, their order, promises and errors as Chromium's own element has them, compared observation by observation (test/polyfill.js). import { Audio } from 'audio/polyfill' leaves the global alone.

CLI

npm i -g audio # or: npx audio …
audio voice.wav trim normalize podcast save clean.mp3
audio clean.mp3 check podcast
#   Apple Podcasts  https://podcasters.apple.com/support/893-audio-requirements
#   ✓ Loudness     -16.44 LUFS  -17 to -15
#   ✓ True peak     -1.18 dBTP  ≤ -1
#   pass

A fail exits 1, so the same line guards a build: Check in CI.

Skill

npx skills add audiojs/audio

MCP

npx add-mcp "npx -y audio --mcp"  # asks which of your agents: Claude Code, Codex, Cursor, Gemini CLI, Kimi Code, Pi, OpenCode, Zed…

Prompt: make ~/Desktop/interview.m4a podcast-ready and tell me the loudness before and after

One agent at a time: claude mcp add audio -- npx -y audio --mcp, qwen mcp add audio npx -y audio --mcp, droid mcp add audio "npx -y audio --mcp".

Playground and agents

npx audio --bridge
# audio bridge on http://127.0.0.1:7777
#
#   key     3e37b8c61acf5b50f010852168f4843d
#   agents  Claude Code, Codex, Pi, Kimi Code
#
#   In the playground's Agent panel, paste the key, then Connect. It stays the same next time.

Open the playground, every edit of it a line of audio code, and connect it to the bridge with the key (once: the bridge keeps it), and its chat runs an agent of yours, which measures, looks at, edits and plays the sound open there, and finds which edit did what: it measures the sound before and after each, and what each took out; each tab keeps its conversations, each with the agent, and its model, picked under the message. The bridge finds Claude Code, Codex, Pi, Gemini CLI, Qwen Code, Kimi Code, OpenCode, Kilo Code, Cline, Goose, Factory Droid, Cursor, Augment, Kiro and Mistral Vibe on PATH; any other that speaks ACP runs by its command line, --agent "my-agent --acp".

Any MCP agent gets the playground's tools (state, measure, look, edit, select, play, check, …), the running bridge found by itself: npx add-mcp "npx -y audio --mcp --playground".

An agent thinks with the model its own settings name: Pi takes local ones for good from ollama launch pi --config; Claude Code takes any Anthropic-compatible endpoint from the bridge's environment:

ANTHROPIC_BASE_URL=http://localhost:11434 ANTHROPIC_AUTH_TOKEN=ollama ANTHROPIC_API_KEY= ANTHROPIC_MODEL=qwen3-coder npx audio --bridge
Models ANTHROPIC_BASE_URL, with the provider's key as ANTHROPIC_AUTH_TOKEN
Ollama http://localhost:11434, token ollama (docs)
Z.ai GLM https://api.z.ai/api/anthropic (docs)
Kimi https://api.moonshot.ai/anthropic (docs)
Qwen https://coding-intl.dashscope.aliyuncs.com/apps/anthropic, Coding Plan (docs)
DeepSeek https://api.deepseek.com/anthropic (docs)

Recipes

Clean up

// master a raw take
let a = audio('raw-take.wav')
a.trim(-30).normalize('podcast').fade(0.3, 0.5)
await a.save('clean.wav')

// full restoration chain via ecosystem plugins (see API › Plugins)
a.gate(-45).dehum().deesser().compressor({ threshold: -18 }).limiter({ ceiling: -1 })

// cut 2:00–2:15, smooth the splice
a.remove({ at: 120, duration: 15 }).fade(0.1, { at: 120 })

// find clipped blocks
let clips = await a.stat('clipping')

// does it pass? each rule of the spec, measured
let { pass, rules } = await a.check('podcast')         // acx, podcast, streaming, broadcast, netflix

Master & deliver

// master a song to a reference track: its tone (mid and side), width and loudness, under -1 dBTP
audio('mix.wav').master(await audio('reference.wav')).save('master.wav')

// audiobook chapter for ACX: RMS -23..-18 dB, peaks under -3 dB, floor under -60 dB, room tone at each end, 44.1 kHz
let ch = audio('chapter-01.wav')
  .highpass(80).omlsa({ gMin: -12 }).compressor({ threshold: -24, ratio: 2.5 })
  .normalize(-20, 'rms', { ceiling: -3.5 })
  .trim().pad(1.5, 2).roomtone()                      // room tone, not digital silence
  .resample(44100)
console.log(await ch.check('acx'))                    // passed 10 of 10 real narrations (.work/pro.md)
await ch.save('chapter-01.mp3', { bitrate: 192 })

// tighten pauses in a talking-head video, then hand the cuts to the video editor
let talk = audio('talk.mp4').shrink(0.3)
await talk.save('talk.m4a')                            // the sound, cut
await talk.save('talk.edl')                            // the same cuts for the picture (Premiere, Resolve)

Compose

// podcast montage
let ep = audio([intro, interview.trim().normalize('podcast'), outro], { crossfade: 0.5 })
await ep.save('episode.mp3')

// voiceover over music
music.gain(-12).mix(voice, { at: 2 })

// ringtone: the chorus + fades
audio('song.mp3').crop({ at: 45, duration: 30 }).fade(0.5, 2).normalize().save('ringtone.mp3')

// split an audiobook into chapters
let [ch1, ch2, ch3] = audio('audiobook.mp3').split(1800, 3600)

// glitch: stutter + reverse
let v = a.clip({ at: 1, duration: 0.25 })
audio([v, v, v, v]).reverse({ at: 0.25, duration: 0.25 })

Analyze

// waveform bars — and progressively, as it decodes
let [mins, peaks] = await a.stat(['min', 'max'], { bins: canvas.width })
a.on('data', ({ delta }) => appendBars(delta.max[0], delta.min[0]))

// features for ML
let mfcc = await a.stat('cepstrum', { bins: 13 })
let [loud, rms] = await a.stat(['loudness', 'rms'])

// notes, chords, key
let notes = await a.stat('notes')    // [{time, duration, freq, midi, note, clarity}]
let chords = await a.stat('chords')  // [{time, duration, label, root, quality, confidence}]
let key = await a.stat('key')        // {tonic, mode, label, confidence}

Record & generate

// mic take
let a = audio()
a.record()
// …later
a.stop()
a.trim().normalize()

// tone — any t => sample function
let tone = audio.from(t => Math.sin(440 * Math.PI * 2 * t), { duration: 2 })

// sonify data
let s = audio.from(t => Math.sin((200 + data[t / 0.2 | 0]) * Math.PI * 2 * t) * 0.5, { duration: data.length * 0.2 })

Automate

a.gain(t => -12 * (0.5 + 0.5 * Math.cos(t * Math.PI * 4)))  // 2Hz tremolo in dB
a.lowpass(t => 400 + 4000 * t)                              // filter sweep
a.pan({ t: [0, 2, 4], v: [-1, 1, -1] })                     // curve L→R→L over 4s, serializable
music.ducker({ key: voice })                                // sidechain (plugin)

Stream & persist

// stream to network — encode/playback during decode
for await (let chunk of audio('2hour-mix.flac').highpass(40)) socket.send(chunk[0].buffer)

// serialize edits, restore later
let json = JSON.stringify(a)     // { source, edits, ... }
let b = audio(JSON.parse(json))  // re-decode + replay edits

API

Create

Method                         Description                                                                                                                        
audio(source, opts?) decode from file, URL, bytes, or a byte stream. Returns instantly — decodes in background; streams decode as they arrive.
audio.from(source, opts?) wrap existing PCM, AudioBuffer, silence, or function. Sync, no I/O.
let a = audio('voice.mp3')                // file path
let b = audio('https://cdn.ex/track.mp3') // URL
let c = audio(inputEl.files[0])           // Blob, File, Response, ArrayBuffer
let d = audio()                           // empty, ready for .push() or .record()
let e = audio([intro, body, outro])       // concat (virtual, no copy)
let f = audio([a, b, c], { crossfade: 2 })  // concat with 2s crossfade
let g = audio(process.stdin)              // byte stream: pipe, socket, fetch body – decodes as it arrives
// opts: { sampleRate, channels, crossfade, curve, storage: 'memory' | 'persistent' | 'auto' }

await a    // await for decode — if you need .duration, full stats etc

let a = audio.from([left, right])                 // Float32Array[] channels
let b = audio.from(3, { channels: 2 })           // 3s silence
let c = audio.from(t => Math.sin(440*TAU*t), { duration: 2 })  // generator
let d = audio.from(audioBuffer)                   // Web Audio AudioBuffer
let e = audio.from(int16arr, { format: 'int16' }) // typed array + format
let f = audio.from(store, { length, channels, sampleRate }) // pages in a store ({ read(i), has(i), write(i, page) }), read as needed

Properties

Property                         Description                                                                                                                        
.duration total seconds, after edits.
.channels channel count.
.sampleRate sample rate.
.length samples per channel.
.currentTime what the speakers play now, in seconds: their latency compensated, smooth; at pause it holds where playback resumes.
.playing true during playback.
.paused true when paused.
.volume 0..1 linear. Settable.
.muted mute, independent of volume. Settable.
.loop settable, mid-playback too: the span repeats, each seam a 10 ms equal-power crossfade.
.playbackRate 0.0625..16, settable mid-playback, click-free. The pitch stays (WSOLA, as a browser's media element plays at a speed) unless .preservesPitch is false: then it glides over ~50 ms and the pitch follows, tape-style. .speed() bakes it.
.preservesPitch true: at a rate other than 1, the pitch kept. false: varispeed, the pitch with the speed. Settable mid-playback.
.ended true when playback reached the end, not after stop().
.seeking true during a seek.
.played promise, resolves when playback sounds.
.recording true during mic recording.
.ready promise, resolves when fully decoded.
.source original source.
.bitDepth stored sample depth of the source: 16, 24, 32 (float); null for lossy or generated audio. Lossless save keeps it.
.pages Float32Array page store.
.stats per-block stats (peak, rms, etc.).
.edits edit list.
.version increments on each edit.

Structure

Method                         Description                                                                                                                        
.trim(threshold?) strip leading/trailing silence (dB, default auto). On a live stream a given threshold streams, holding a silent tail until sound resumes; the automatic one reads the whole input, so it waits for the end.
.shrink(gap?, threshold?) shorten silent pauses to gap seconds (default 0.3); 0 removes them. Streams with a given threshold, like trim.
≡ FFmpeg silenceremove, Audacity truncate-silence
.crop({at, duration}) keep range, discard rest.
.remove(at, duration, crossfade?) delete range, close gap. crossfade ('10ms') makes the splice an equal-power crossfade centered on the cut; the length stays the same.
.insert(source, at?, crossfade?) insert audio (default: at end), or a number of seconds of silence; crossfade fades both seams.
.copy({at?, duration?}) copy range (default: all) to this instance's clipboard.
.cut({at?, duration?, crossfade?}) copy, then remove.
.paste({at?, crossfade?}) insert the clipboard (default: at end).
.move({at, duration, to, crossfade?}) slide a range to to, over what is there; silence where it was, the length kept (past the end, extended). crossfade crossfades each edge, centered.
≡ a DAW's clip moved in slip mode
.clip({at, duration}) zero-copy excerpt as a new instance.
.split(...offsets) zero-copy excerpts between timestamps.
.pad(before, after?) silence at edges (seconds).
.repeat(n) repeat n times.
.reverse({at?, duration?}) reverse audio or range.
.speed(rate) changes pitch and duration together.
.stretch(factor, {voice?}) changes duration, keeps pitch (phase-locked vocoder). A t => f or {t, v} factor slides the tempo; duration becomes ∫factor dt. A range {at, duration} comes out round(round(duration · sampleRate) · factor) samples long, to the sample; the audio around it as it was. { voice: true } keeps a voice's pulse shape and consonants, which the vocoder makes distant: shortened, its waveform copied a segment at a time (WSOLA, @audio/stretch-wsola); slowed, the vocoder's frames restarted from the waveform where it fits (PVSOLA, @audio/stretch-pvsola), so breath and reverberation are not repeated into a flanger. One voice, not chords.
≡ Logic Flex Time Monophonic, Ableton Tones
.warp(markers) move moments in time: [[from, to], …] in seconds. Between markers the audio stretches to fit, pitch kept; start and end stay.
≡ Logic Flex Time, Ableton warp markers
.pitch(semitones, {voice?}) changes pitch, keeps duration. Semitones may be a curve {t, v} (seconds → semitones, straight between points, flat past the ends, as the gain line's) or t => semitones; where it is zero the audio is as it was. { voice: true } re-spaces a voice's own glottal cycles (TD-PSOLA, the optional @audio/tune-curve): formants and consonants kept, one voice.
≡ Melodyne pitch drawing
.intonation(factor) a voice's rises and falls wider or flatter about its median pitch: 1 as it was, 0 a monotone, 2 twice as wide. Its own cycles re-spaced (as pitch({ voice: true })): timing, formants and consonants kept, one voice.
≡ Melodyne pitch modulation, Praat's pitch range factor
.formant(semitones) moves the formants (the spectral envelope: a voice's vowels, the size of its head), keeps the pitch. A number, a curve {t, v} or t => semitones. Any sound; with pitch(), a voice kept its own or made another's.
≡ Melodyne formant tool, Praat's formant shift ratio
.remix(channels) channel count (down per ITU-R BS.775: 7.1 → 5.1, stereo, mono), or a map: [1, 0] swaps L/R, null a silent channel.

Every op takes a trailing {at, duration, channel} range, except channel-changing remix and crossover, and every op that keeps the timeline a mix: how much of its output is heard, the rest its input, 0 to 1 or a curve {t, v} over time ({ mix: { t: [11.99, 12, 14, 14.01], v: [1, 0, 0, 1] } } turns it off from 12 s to 14 s, as RX's Restore Selection; an effect's own mix is its own). Times are seconds or strings ('1:30', '2m'); negative counts from the end. FFmpeg's short names work wherever the long ones do: d for duration, xfade for crossfade (a.remove({ at: 1, d: 0.5, xfade: 0.01 })).

a.trim(-30)                               // strip silence below -30dB
a.remove({ at: '2m', duration: 15 })      // delete 2:00–2:15, close gap
a.remove(12.3, 0.4, '10ms')               // cut a breath, crossfaded: no click
a.insert(intro, { at: 0 })                // prepend; .insert(3) appends 3s silence
a.copy(60, 30).paste(120)                 // duplicate the chorus at 2:00
a.cut(2, 1).paste(5)                      // move 2s–3s to 5s of the shortened timeline
a.move({ at: 2, duration: 1, to: 5 })     // slide 2s–3s over 5s–6s, silence left at 2s–3s
let [pt1, pt2] = a.split('30m')           // zero-copy parts
let hook = a.clip({ at: 60, duration: 30 })  // zero-copy excerpt
a.stretch(1.1)                            // 10% longer, same pitch
a.warp([[1, 1], [2, 2.4], [3, 3]])        // the hit at 2s lands at 2.4s; 1s–3s keeps its length
a.pitch(-2)                               // 2 semitones down, same tempo
a.pitch({ t: [1, 1.2, 2, 2.2], v: [0, 3, 3, 0] }, { voice: true })  // a note drawn 3 semitones up, formants kept
a.intonation(1.5, { at: 2, d: 3 })        // 2s–5s: every rise and fall half as wide again
a.formant(-2)                             // a larger, darker voice on the same notes
a.remix([0, 0])                           // L→both; .remix(1) for mono

Process

Method                         Description                                                                                                                        
.gain(dB, opts?) { unit: 'linear' } takes a multiplier.
.fade(in, out?, curve?) curves 'linear' 'exp' 'log' 'cos', as functions in audio.op('fade').curves. {start, end} levels 0..1 fade between any levels (a duck); {mid} skews the half-amplitude point.
≡ Audacity adjustable-fade
.normalize(target?, mode?) remove DC, normalize. Loudness targets hold a true-peak ceiling, -1 dBTP by default: a lookahead limiter, then the loudness it took made back up. Presets per Apple Podcasts, Spotify, EBU R 128 (ITU-R BS.1770-4):
'podcast' -16 LUFS
'streaming' -14 LUFS
'broadcast' -23 LUFS
-18, 'lufs' any loudness; -3 peak dB; no arg: peak 0 dBFS; 'rms' mode
an audio instance: its integrated loudness
{ ceiling: -2 } dBTP, false off
{ dc: false } keep DC
{ adaptive: true } on a live stream, start at once: the gain follows what it has heard, the ceiling (the target itself in peak mode) guards what it hasn't. Without it, one gain for the whole selection: a live stream waits for its end.
≡ FFmpeg loudnorm
.roomtone(threshold?) fill digital silence (≥ 10 ms under -90 dBFS: edited-out pauses, pad()) with the recording's own room tone: its quiet stretches that hold still for 0.3 s, 20 dB under the program. .trim().pad(1.5, 2).roomtone() gives an audiobook chapter its room tone at each end (ACX rejects digital silence). Where every pause was gated or cut, there is no room to take, and nothing changes.
≡ iZotope RX Ambience Match
.mix(source, at?, gain?) overlay at at seconds, source level gain dB.
≡ FFmpeg amix weights
.crossfade(source, duration?, curve?) append with overlap, default 0.5s. 'cos' (default) suits similar material; 'equal' (equal-power) keeps loudness across unrelated tracks; each in audio.op('crossfade').curves. With no source, .crossfade({ at, duration }) crossfades across the range, as an editor crossfades a selection: the audio before it fades into the audio after it, and the range goes.
≡ FFmpeg acrossfade
.pan(value, opts?) −1 left, 0 center, 1 right.
.write(data, {at?}) overwrite from at with raw PCM or another sound, as a tape records over what is there; what runs past the end extends it.
.transform(fn) inline (input, output, ctx) => void. Not serialized.
a.gain(-3)                                // reduce 3dB
a.gain(6, { at: 10, duration: 5 })        // boost range
a.gain(t => -12 * Math.cos(t * TAU))      // automate over time
a.fade(0.5, -2, 'exp')                    // 0.5s in, 2s exp fade-out
a.normalize('podcast')                    // -16 LUFS, -1 dBTP
a.normalize(-27, 'lufs', { ceiling: -2 }) // Netflix
a.mix(voice, { at: 2 })                   // overlay at 2s
a.mix(bed, 0, -18)                        // music bed, 18 dB under
a.crossfade(next, 2)                      // 2s crossfade into next
a.crossfade(song2, 4, 'equal')            // equal-power, for unrelated tracks
a.pan(-0.3, { at: 10, duration: 5 })      // pan left for range

Filter

Method                         Description                                                                                                                        
.highpass(freq, order?), .lowpass(freq, order?) Butterworth pass filter; even integer order ≥ 2: 2 (12 dB/oct, default), 4 (24), 6, 8, … Other orders are rejected.
.bandpass(freq, Q?), .notch(freq, Q?) band-pass / notch.
.allpass(freq, Q?) phase shift, unity magnitude.
.phase(angle?) every frequency's phase turned by angle degrees (a Hilbert transform; a number, a curve {t, v} or t => degrees): magnitudes, RMS and loudness as they were, and -angle turns it back. 180 inverts polarity, sample for sample ({ channel: 1 }: one channel's). Unset, the angle over time that lowers the peaks most, the channels together, never raising one: a voice stands more evenly about zero: its peak 0.7 and 0.8 dB lower on average on VoiceBank's two test speakers (30 sentences each), up to 2.9; where two voices take turns, each its own angle. Streams 0.28 s behind.
≡ iZotope RX Phase
.lowshelf(freq, dB), .highshelf(freq, dB) shelf EQ.
.eq(freq, gain, Q?) parametric EQ.
.filter(type, ...params) by type name, or a custom filter function.

All biquads.

a.highpass(80).lowshelf(200, -3)          // rumble + mud
a.eq(3000, 2, 1.5).highshelf(8000, 3)     // presence + air
a.notch(50)                               // remove hum
a.allpass(1000)                           // phase shift at 1kHz
a.phase()                                 // the peaks lowered, nothing heard
a.phase(180, { channel: 1 })              // the right channel's polarity
a.filter(customFn, { cutoff: 2000 })      // custom filter function

Effect

Method                         Description                                                                                                                        
.vocals(mode?, {model?}) mid/side: 'isolate' (default) keeps center, 'remove' keeps sides. model separates with a trained model instead, 'scnet-large' (SCNet, MIT weights, the highest SDR here), 'umxhq' (Open-Unmix, MIT weights) or 'htdemucs' (Hybrid Transformer Demucs, weights for research only), through the optional @audio/neural-separate; 'remove' then subtracts the model's vocals. Weights are exported locally (how) or served from weights.
≡ SoX oops; SCNet, Demucs, Open-Unmix
.rebalance(vocals?, bass?, drums?, other?, {model?}) a song's vocals, bass, drums and the rest, each at its own level, dB (−Infinity mutes): a separation model splits the input into the four stems and adds each one's change to it, so 0 dB each leaves the input sample for sample and what the model gives no stem stays as it was. model 'scnet-large' (default: SCNet-large, MIT weights, 169 MB), 'scnet' (43 MB, a third of the time), 'htdemucs' (weights for research only) or 'umxhq', through the optional @audio/neural-separate, its weights exported locally (how) or served from weights; the input is separated once, gains moved read the stems back. MUSDB18 test previews, BSSEval v4 SDR of each stem soloed, vocals · bass · drums · other: 10.75 · 8.17 · 10.30 · 6.94, iZotope RX 12 Music Rebalance at Best 10.89 · 9.66 · 9.85 · 6.45, ahead on 33, 32, 38, 32 of the 50 songs; vocals 6 dB up 20.2 dB against 19.5, drums 6 dB down 21.9 against 20.9 (bench/rx/separate.mjs).
≡ iZotope RX Music Rebalance
.scene(dialogue?, music?, effects?, {model?}) a soundtrack's dialogue, music and effects, each at its own level, dB (−Infinity mutes): a separation model splits the input into the three stems and adds each one's change to it, so 0 dB each leaves the input sample for sample. model 'mrx' (default: MRX, Petermann et al., ICASSP 2022, MERL's MIT weights, 122 MB, a fifth of real time) or 'tiger' (TIGER, ICLR 2025, Apache-2.0 weights, 29 MB, about 50 times the time: on 30 of the test clips 12.7 · 10.2 · 8.1 dB against MRX's 10.9 · 5.2 · 5.7, the dialogue 6 dB up 21.0 against 18.7), through the optional @audio/neural-separate, its weights exported locally (how) or served from weights; the input is separated once, gains moved read the stems back. Divide and Remaster v3's English test set, 150 clips of 60 s, SNR median, dialogue · music · effects: 11.8 · 5.3 · 6.0 dB (the dialogue as deepfilter(0) takes it: 10.5; its authors' Bandit v2, weights CC BY-SA: 15.6 · 10.4 · 9.9); the dialogue 6 dB up 18.9 dB from the true remix, against 7.2 left as it is (bench/rx/scene.mjs). RX 12 Scene Rebalance runs only inside Pro Tools (AAX): not measured.
≡ iZotope RX Scene Rebalance
.dither(bits?, {shape?}) TPDF, default 16-bit. shape: true adds 2nd-order noise shaping: quantization noise moves above ~Nyquist/2, audibly quieter.
.crossfeed(freq?, level?) headphone crossfeed, default 700 Hz, 0.3.
≡ SoX earwax, bs2b
.convolve(ir, mix?, {normalize?, tail?}) the sound through an impulse response: a room, a plate, a cabinet, a microphone as captured. ir a file or URL, an audio instance or channels of samples, resampled to the sound's rate; channel c through its channel c, wrapping. No latency (Gardner's partitioned convolution: a direct head, FFT partitions behind it), its decay rendered past the end (tail: false, none). mix 0 to 1 (1), normalize: true the IR at unit energy.
≡ Pedalboard Convolution, FFmpeg afir
.resample(rate, {type?}) upsampling defaults to linear, downsampling to anti-aliased windowed sinc, its taps widening with the ratio. type: 'sinc' or 'linear' forces one.
.crossover(...freqs) N split frequencies → N+1 bands × channels, band-major. Linkwitz-Riley 4th order; bands sum back flat.
≡ FFmpeg acrossover
.match(ref, amount?) match EQ: up to 8 parametric bands fit to the reference/source spectrum ratio. Tone only; loudness stays with normalize. Streams {lookahead} s behind (10), refitting as it hears more. { midside: true } matches a stereo pair's mid and side apart, and the side level to the reference's width.
≡ iZotope Ozone Match EQ
.master(ref, opts?) master to a reference track: match in mid and side, then normalize to the reference's integrated loudness under -1 dBTP ({ ceiling }).
≡ Matchering
.auto(type?, {intensity?, targetLufs?, ceiling?}) the take repaired and finished by what it carries (@audio/chain): measured, then each repair on its own evidence (declip, declick, deplosive, dehum, denoise, dereverb, deesser), and only where measured need calls for them an EQ (a voice past the spread clean voices keep about the speech target), glue (a loudness range over 12 LU), and loudness at the type's target ('speech', default, −16 LUFS; 'music' −14; 'voice-music'), its true peak under ceiling dB (−1) by up to 6 dB × intensity of true-peak limiting where the peaks need it: a take that needs nothing comes back as it was but for its level. A bed under speech goes to DeepFilterNet3 where @audio/neural-denoise is installed (deepfilter()'s model, the noise 60 dB × intensity down, the recipe naming it), else to OM-LSA; under 'speech' the voice is what is kept. intensity 0–2 (1) scales how hard; 0 stops the gain at the ceiling. On takes with several defects made on clean recordings (bench/rx/assistant.mjs), its repairs against iZotope RX 12 Repair Assistant tuned: speech PESQ 2.99 against 2.56, SI-SDR 18.4 dB against 13.0, DNSMOS OVRL 3.12 against 3.12, music ODG −1.69 against −2.23; clean speech 4.46 against 3.81; steady beds level (PESQ 2.97 against 2.94), RX ahead on DEMAND's (2.53 against 2.90: four of nine read as no bed). Without the neural package speech 2.79, beds 2.39 and 2.07. auto() whole: 89 % of speech and 98 % of music within 1 LU of the target, true peak at most −1.00 dBTP; the limiting costs a clean mix ODG 0.30 against the repairs (−0.22, input 0.21), a clean voice nothing (PESQ 4.45).
≡ iZotope RX Repair Assistant, Dolby.io Media Enhance
.spectral(band?, gain?, {at, duration}) gain on a time × frequency region, band = [lo, hi] Hz; default removes it.
≡ Audacity spectral edit, FFmpeg afftfilt
.repair(band?, {at, duration, method?, window?}) rebuild a damaged range (dropout, beep, click burst) from its surroundings. method 'auto' (default) transplants the passage that joins seamlessly, searched in the window s (10) before the range; failing that, AR interpolation up to 70 ms, a sinusoidal bridge beyond. 'ar', 'sinusoidal', 'similarity', 'spectral' force one.
≡ iZotope RX Spectral Repair
.declick(threshold?, longest?, {at?, duration?}) remove clicks: a vinyl tick, a bad splice, a digital glitch, a mouth click on a voice. Each stands threshold times (8) out of the AR prediction error around it; how far it reaches and whether it rings on are what explains the sound best, and it is rebuilt as the sound's most likely value under it from 46 ms either side; a pulse with its like 2.5–15 ms away (a voice, a plucked string), the sound's own excitation (a plosive, an attack) or a burst over longest ms (6) is sound, left alone. Given { at, duration }, the clicks there: looked for only there, none passed over for its likes or its length. Nothing a click doesn't reach changes. Streams about 0.9 s behind. Clicks at 5× the sound around them: ticks 34 dB down, glitches 46, mouth clicks 28, more than iZotope RX 12 De-click and Mouth De-click take off in every kind and size measured (@audio/denoise-declick's README).
≡ iZotope RX De-click, Audacity Click Removal
.denoise(reduction?, threshold?, {noise}) remove a noise that holds still (hiss, hum and buzz, a fan, room tone, tape), learned where it plays alone: noise is that { at, duration } of the op's input, or several, or a print saved from stat('print'). It goes reduction dB down (12) everywhere, or in the op's own { at, duration }, or only in a band [low, high] Hz, the rest as it was; what stays is the same noise, quieter, without musical tones. threshold (dB) raises the print: more of the quiet counts as noise. OM-LSA on the held noise (@audio/denoise-omlsa), each channel its own print; a live source renders once the range has arrived. VoiceBank+DEMAND PESQ, the noise learned from the half second before each speaker starts: noisy 1.97, omlsa() 2.40, denoise() 2.48. For noise that moves: omlsa(), deepfilter().
≡ iZotope RX Spectral De-noise (Learn), Adobe Audition Noise Reduction (noise print), Audacity Noise Reduction
.deepfilter(limit?, {floor?, music?, weights?, device?}), .rnnoise(limit?) neural speech denoising through the optional @audio/neural-denoise: it also removes noise that moves (keys, traffic, a busy room). deepfilter runs DeepFilterNet3, its 8 MB model downloaded once, over the whole input before rendering; rnnoise streams RNNoise, weights in the package, 30 ms behind. limit is how far the noise goes down, in dB: deepfilter's 18 is the most before the voice itself sounds filtered (DNSMOS SIG holds to 18 and falls past it), so room tone stays, and noise closer to the voice than floor dB (40) goes down to that floor, so noise as loud as the voice is taken, not left 18 dB under it; rnnoise's 16 is the most before it cuts into the voice: RNNoise turns down word ends even with no noise at all (10 dB at 16, 31 unlimited); 0 lifts it. deepfilter keeps held sung notes, which the model alone takes for noise (VocalSet: 1.9 dB down, not 39), and hears a 16 kHz file's or a codec's empty top band as the noise floor the model trained with (16 kHz VoiceBank+DEMAND PESQ 2.83, not 2.72). VoiceBank+DEMAND PESQ: noisy 1.97, wiener() 2.19, rnnoise() 2.49, deepfilter() 3.05, deepfilter(0) 3.15, iZotope RX 12 Dialogue Isolate 2.72; in noise, babble and music beds 5 to −5 dB under the voice and in rooms, deepfilter() 1.99 against RX's 1.86 (bench/rx/isolate.mjs). Music alone passes untouched (music: 'enhance' enhances it too).
≡ DeepFilterNet, RNNoise, iZotope RX Dialogue Isolate
.derustle(reduction?, {ambience?}) a clip-on mic's clothing rustle taken off the voice, the room's steady tone kept: deepfilter's model (the two share its run on the same input), what it took mixed back as deepfilter mixes it, rustle down to reduction dB under the voice (45), the room's tone 12 dB down (ambience: false takes it all). Settings chosen on clothing rustle under 10 to 37 s talkers (Freesound's CC0 and CC BY clothing foley and mic handling: no lavalier rustle is openly licensed). Test, 55 talkers, rustle as loud as the voice: PESQ 1.21 → 2.13, SI-SDR −0.3 → 13.5 dB; the room alone 4.37, a clean take 4.61 (deepfilter() 2.12, 4.31, 4.59; iZotope RX 12 Dialogue Isolate, RX's hostable dialogue separator, 2.06, 2.96, 3.96) (bench/rx/derustle.mjs). RX 12 De-rustle runs only inside Pro Tools (AAX): not measured.
≡ iZotope RX De-rustle
.deconstruct(tonal?, noise?, transient?, {separation?}) the tonal, noisy and transient parts, each at its own level, dB: median filters over time and over frequency (Fitzgerald 2010), the noise what is neither by separation (2; Driedger, Müller & Disch 2014), soft masks the channels share. The parts add back to the input; 0 dB each leaves it sample for sample. A chord lands 100 % in the tonal part, clicks in the transient, a hiss 68 % in the noise (1 % in each other).
≡ iZotope RX Deconstruct
.azimuth(delay?) a stereo pair lined up in time and polarity, the second channel moved onto the first. Unset, the delay is measured every second (generalized cross-correlation with Knapp & Carter's 1976 weighting, refined between samples: an exact fractional shift comes back to 0.0001 samples) and followed as it wanders; a channel wired backwards turned back; a pair already in line, as a mix's stereo image is, left sample for sample. delay ms sets it by hand.
≡ iZotope RX Azimuth
.dewow(mode?) wow and flutter: the speed of the disc or tape measured over time and the sound read back at its inverse; from the music's own notes (a disc's once-a-turn wow, 'partial'), a pilot or calibration tone ('reference', refFreq), or one voice's pitch ('pitch'). Nothing proven, nothing changed: a recording without wow comes back sample for sample (@audio/denoise-dewow).
≡ iZotope RX Wow & Flutter
.desqueak(squeak?, {pick?, amp?}) a guitar's string squeaks taken down, where a finger slides along a wound string: told from the notes by what a note is not (partials that hold their bins, read on 93 ms frames; a pluck's low partials or click), held 45 ms or more, each bin down to what it holds around the squeak, squeak dB at most (30); the notes ringing under it keep their partials, and the take outside the squeaks stays sample for sample. pick dB softens each pluck's click to the level its note rings at 15 ms on (-9 restores attacks made 6–12 dB harsher); amp dB takes a steady hiss and buzz down (OM-LSA on a print from the quietest frames, a note's peaks kept). GuitarSet test takes against iZotope RX 12 Guitar De-noise: squeaks added, 3.3 dB of their error gone against 2.7, SNR 23.2 dB against 20.5, 2.4 % of a clean take's samples moved against 12.9 %; harsh picks 8.9 dB against 3.3 tuned; squeaks as recorded 12.8 dB median against 13.5, half the plucks and clicks RX touches (@audio/denoise-desqueak, bench/rx/guitar.mjs).
≡ iZotope RX Guitar De-noise
.codec(format?, bitrate?, {quality?}) the sound as a lossy codec gives it back: 'mp3' (default), 'aac', 'opus' or 'vorbis' at bitrate kbps (128), 'mp3' VBR at quality (LAME's -V: 0 best, under 10), 'gsm' (GSM 06.10 full rate, 13 kbps at 8 kHz: a 1990s phone line), encoded and decoded in place. Each codec's delay undone (its own gapless information, then checked by correlation), so it lines up with the input within 0.03 of a sample (at 44.1 and 48 kHz): A/B it, or hear what it takes out. A codec moves the peaks: masters normalized to -14 LUFS under -1 dBTP came back up to 0.5 dB over it through Opus at 160 kbps, 0.3 through Vorbis, under 0.1 through MP3 320 and AAC 256 (7 s excerpts of 8 MUSDB18 mixes); .codec('opus', 160).stat('truepeak') measures it.
≡ iZotope RX Streaming Preview, Pedalboard MP3Compressor, GSMFullRateCompressor
a.vocals()                                // isolate center-panned vocals
a.vocals('remove')                        // remove vocals (karaoke)
a.vocals({ model: 'umxhq' })              // vocals by a separation model
a.rebalance({ vocals: 6, drums: -3 })     // the vocals up, the drums down
a.rebalance(-Infinity)                    // a karaoke track
a.scene(6)                                // a film's dialogue 6 dB up over its music and effects
a.scene(-Infinity)                        // a music and effects track
a.dither(16)                              // TPDF dither to 16-bit
a.dither(16, { shape: true })             // noise-shaped
a.crossfeed()                             // headphone crossfeed
a.convolve('hall.wav', 0.3)               // 30 % of a captured hall
a.resample(48000)                         // resample to 48kHz (linear)
a.resample(96000, { type: 'sinc' })       // high-quality windowed-sinc
a.match(reference, 0.7)                   // 70% of the way to its tone
a.spectral([1000, 4000], -30, { at: 2.1, duration: 0.3 })  // a cough
a.repair({ at: 1.2, duration: 0.05 })     // a dropout
a.repair({ at: 42, duration: 1 })         // a lost second of music: the passage that fits
a.declick()                               // every click
a.declick({ at: 12.31, duration: 0.02 })  // the clicks seen there
a.denoise({ noise: { at: 1.2, duration: 0.5 } })  // hiss learned from a pause, 12 dB down everywhere
a.deepfilter()                            // speech out of noise: room tone 18 dB down and kept, loud noise to 40 under the voice
a.rnnoise()                               // the same, streaming
a.derustle()                              // a clip-on mic's clothing rustle, the room kept
a.deconstruct(0, -12)                     // the hiss 12 dB down, tones and attacks as they were
a.azimuth()                               // a tape's two tracks back in line
a.dewow({ mode: 'reference', refFreq: 1000 })  // a transfer's wow, read from its calibration tone
a.desqueak()                              // a guitar's string squeaks
a.desqueak({ pick: -9, amp: -20 })        // and its harsh picks, and the amp's hiss and buzz
a.codec('aac', 256)                       // what a 256 kbps AAC stream does to it
a.codec('gsm')                            // a phone call

I/O

Method                         Description                                                                                                                        
await .read(opts?) rendered PCM. { format, channel } to convert. A source still arriving is waited for: a range until it has arrived, all of it until the end (an endless stream: read ranges, or stream()).
await .save(path, opts?) encode + write, format from extension. Lossless keeps the source depth; { bitDepth, bitrate, quality, codec } set the encoder; m4a and mp3 write markers as chapters. Output streams as it encodes, headers patched with their totals at the end (a pipe keeps them "unknown"); m4a from a live source is fragmented. A video source saved to .mp4/.mov keeps its picture: only the audio track changes. .edl, .otio, .fcpxml write the cuts.
await .encode(format?, opts?) encode to Uint8Array; 'edl', 'otio', 'fcpxml': the cuts, as UTF-8.
await .cuts(format?, opts?) the edits as a cut list for a video editor: 'edl' (CMX 3600: Premiere, Resolve, Avid), 'fcpxml' (Final Cut Pro, Resolve), 'otio' (OpenTimelineIO); none gives { fps, clips: [{ at, duration, from, rate, source }] }. Cuts, moves, gaps and inserted files place the clips; speed, stretch and warp set their rate; markers go along. Processing is not in the list: lay the processed audio under the picture. Cuts land on the video track's frames (MP4/MOV), else { fps } (30), each within half a frame of the sound; its timecode track starts the source times, drop-frame as it counts ({ dropFrame }). The CLI's save cuts.edl writes one.
.clone() independent edits, shared pages.
.push(data, format?) feed PCM into a pushable instance; .stop() finalizes.
let pcm = await a.read()                              // Float32Array[]
let raw = await a.read({ format: 'int16', channel: 0 })
for await (let block of a) send(block)                 // async-iterable over blocks
await a.save('out.mp3')                                // format from extension
await a.save('book.mp3', { bitrate: 192 })             // ACX: 192 kbps CBR
await a.save('master.wav', { bitDepth: 24 })           // 24-bit
await a.save('talk.mp4')                               // video in, video out
let bytes = await a.encode('flac')                     // Uint8Array
let b = a.clone()                                      // independent copy, shared pages

let src = audio()                                      // pushable source
src.push(buf, 'int16')                                 // feed PCM
src.stop()                                             // finalize

Playback / Recording

Method                         Description                                                                                                                        
.play(opts?) { at, duration, loop, volume, rate, paused, device }. at defaults to currentTime (the start once ended); playing already, it jumps there without a gap. device: an output by id or name (a page plays it on a context of its own, sunk there).
.play({ from: b }) take over b's playback where it is (its span, loop, volume, rate, pause), crossfaded, no gap; b stops.
.pause(), .resume(), .seek(t), .stop() each ramps over 5 ms, none clicks; seek crossfades, and in a loop stays in its span. stop() also ends recording.
.record(opts?) mic. { device, sampleRate, channels, monitor }. device: an input by id or name. monitor: the take through its edits as it comes in, to the default output (true) or one by id or name, about 50 ms behind (the input's read, one block, the output's ring; and the edits' own lookahead): input monitoring, as a DAW's through a track's inserts.
≡ Pedalboard AudioStream
audio.devices() the inputs and outputs: { input: [{ id, name, default }], output: [...] }. Node: the platform's own ids (CoreAudio's UID, WASAPI's endpoint id, ALSA's name); a page: enumerateDevices(), names once the microphone is allowed.
audio.context the page's one AudioContext, which playback uses: made on first use, resumed by the first gesture; set your own before playing.

Playback renders up to 2 s ahead into an AudioWorklet on audio.context (Node: @audio/speaker), so a busy main thread doesn't stop it, and sounds within milliseconds of play() (the device's own latency aside). An edit to the playing instance is heard ~50 ms later where it happens: the audio rendered ahead gives way, crossfaded. A source still arriving (decoding, pushed) plays what has come and goes on as more comes. Any channel count plays as it is; the device downmixes.

a.play({ at: 30, duration: 10 })          // play 30s–40s
await a.played                            // wait for sound
a.volume = 0.5; a.loop = true             // live adjustments
a.muted = true                            // mute without changing volume
a.playbackRate = 1.5                      // faster, the pitch kept
a.preservesPitch = false                  // the pitch follows the speed, as a tape's
a.pause(); a.seek(60); a.resume()         // jump to 1:00
a.highpass(80)                            // an edit while playing: heard where it happens
b.play({ from: a })                       // b takes over at the same place, crossfaded
b.stop()                                  // end playback or recording

await audio.context.audioWorklet.addModule('./scrub.js')  // your own nodes, on the same context
let scrub = new AudioWorkletNode(audio.context, 'scrub')

let mic = audio()
mic.record({ sampleRate: 16000, channels: 1 })
mic.stop()

await audio.devices()                     // { input: [{ id, name, default }, …], output: […] }
let take = audio()
take.compressor().plate(0.2)              // edits, heard as they come in
take.record({ device: 'USB', monitor: 'Headphones' })   // by id, name, or part of a name

Metering

Method                         Description                                                                                                                        
.meter(what, cb?) live per-block stats of what plays, delivered as it is heard: rms, peak, ms, min, max, dc, clipping, spectrum, or your own. Without cb, read .value. Returns { value, stop() }. The same on a worker facade: measured in the worker, delivered on the page.
Option                         Description                                                                                                                        
type stat name, array of names, or omit for all block stats.
channel n for one channel, [n, m] per-channel, or omit for scalar avg (mirrors a.stat()).
smoothing one-pole EMA time constant τ, in seconds.
hold peak-hold decay τ, in seconds.
bins, fMin, fMax spectrum resolution and range (when type: 'spectrum').
a.meter('rms', v => draw(v))                                       // scalar avg across channels
a.meter(['rms', 'peak'], v => draw(v))                             // { rms, peak }
a.meter({ type: 'rms', channel: [0, 1] }, v => draw(v))            // [L, R]
a.meter({ type: 'spectrum', bins: 64, smoothing: 0.15 }, drawFFT)  // Float32Array of mel bins
a.meter({}, ({ delta, offset }) => draw(delta))                    // no type → all block stats

let m = a.meter({ type: 'rms' })                                   // pull form
requestAnimationFrame(function tick() { draw(m.value); requestAnimationFrame(tick) })
m.stop()                                                           // release

Analysis

Method                         Description                                                                                                                        
await .stat(name, opts?) one value; with { bins: n } a Float32Array, the value over each of n spans of the range: where, not only how much (lists, a key and spectra come whole; bins sizes spectrum and cepstrum); an array of names gives an array. { channel: n } one channel, [n, m] per channel; {at, duration} sub-range.
await .detect(opts?) { bpm, confidence, beats, onsets } in one pass; { channel } as in stat.
await .check(spec) pass or fail against a delivery spec: { pass, rules: [{ name, value, unit, min, max, pass }] }. 'acx' (RMS, peak, noise floor, room tone, 44.1 kHz), 'podcast' (Apple: -16 LUFS ±1, ≤ -1 dBTP), 'streaming' (Spotify: plays at -14 LUFS, ≤ -1 dBTP), 'broadcast' (EBU R 128: -23 ±0.2 LUFS, ≤ -1 dBTP), 'netflix' (dialog -27 ±2 LUFS, ≤ -2 dBTP). Each limit cites its source in fn/check.js.
Stat                         Description                                                                                                                        
'db' peak amplitude in dBFS.
'rms' RMS amplitude, linear (the CLI prints dBFS).
'noisefloor' RMS of the quietest 0.4 s, dB: the room between words (ACX Check's measure, sample-exact).
'print' the noise print of a range, as denoise({ noise }) takes it: dB in 1025 bands 23.4375 Hz apart, 0 to 24 kHz (white noise of RMS 0.01 prints −40).
'peak' max(|min|, |max|), linear.
'loudness' integrated LUFS (ITU-R BS.1770-4; surround channels weighted, LFE excluded).
'momentary', 'shortterm' maximum 400 ms / 3 s loudness, LUFS (EBU Tech 3341).
'dialog' loudness of the speech only, LUFS: speech found automatically (AES TD1008 dialog loudness).
'dc' DC offset.
'clipping' clipped samples, at 16-bit full scale (±32767/32768) or beyond (scalar: timestamps, binned: counts).
'silence' silent ranges as {at, duration}.
'crest' peak/RMS in dB. Sine ≈ 3dB, square ≈ 0dB.
'centroid' spectral centroid in Hz (brightness).
'flatness' spectral flatness: 0 tonal, 1 noise.
'correlation' L/R phase correlation, −1 to +1. Mono returns 1.
'max', 'min' peak envelope per bin, for waveforms.
'spectrum' mel spectrum in dB (A-weighted); of several channels, their mean power.
'cepstrum' MFCCs.
'bpm' tempo.
'beats', 'onsets' timestamps as Float64Array (seconds).
'hits' where the level jumps, up (a strike) or down (a stop), as Float64Array (seconds): each at its attack's zero crossing, the sample to cut at, where 'onsets' reads 23 ms blocks.
'notes' [{time, duration, freq, midi, note, clarity}] (pYIN + Tony note HMM); with robust: true, the same through noise and rooms (a neural pYIN stage 1); with poly: true, polyphonic [{time, duration, freq, midi, note, velocity, bends}] (Basic Pitch).
'chords' [{time, duration, label, root, quality, bass, confidence}]: Chordino on NNLS chroma (Mauch & Dixon 2010), matched to the reference plugin; labels like 'Am', 'G7', 'C/E', 'N'.
'key' {tonic, mode, label, confidence} (Krumhansl-Schmuckler).
'voicing' share of the range voiced, 0 to 1: frames of pYIN's pitch curve with a pitch, every 10 ms.
'hnr' harmonics-to-noise ratio, dB, over the voiced frames (Boersma 1993; matches Praat's To Harmonicity (ac) frame by frame); null where none is voiced.
'harmonic' level of the periodic part, dB: the periodic share of each frame's power (Boersma's r) times the power. Unlike RMS, unmoved by noise taken away: what an edit left of a voice.
'similar' where else it sounds like { at, duration } (a cough, a click, a beep): [{ at, duration, score }], the range's log-mel patch, each band over its median, slid along the whole by Pearson's correlation, from threshold (0.7); band [low, high] Hz compares those frequencies alone. Also a.similar(range, opts).
≡ iZotope RX Find Similar

Opts: bpm, beats, onsets take { minBpm, maxBpm, delta, frameSize, hopSize }; notes takes { minFreq, maxFreq, frameSize, hopSize, minDuration }; chords, key take { frameSize, hopSize, tuning } (frames of 16384 samples at 44.1 kHz, 0.34 to 0.51 s at other rates, every eighth of a frame; concert A read from the audio unless tuning in Hz is given); chords also boostN (no-chord bias, 0.1); key also method: 'nnls' | 'pcp'. chords needs @audio/mir-nnls-chroma and @audio/mir-chordino, key needs @audio/mir-nnls-chroma: GPL-2.0-or-later translations of the reference plugins, installed by choice (npm i @audio/mir-nnls-chroma @audio/mir-chordino); key with method: 'pcp' needs only the MIT @audio/mir-chroma and @audio/mir-key, installed with audio unless optional dependencies are skipped. notes with robust: true needs @audio/neural-pitch (weights inside): a network's pitch candidates in place of YIN's keep the notes where YIN loses them (Vocadito onsets F 0.76 against 0.53 at 0 dB SNR) and trail it slightly on clean audio, so YIN stays the default. notes with poly: true takes { minFreq, maxFreq, minDuration, onsetThreshold, frameThreshold } and needs @audio/neural-transcribe, whose model downloads on first use; bends are cents from the note's pitch per 11.6 ms frame, in 33.3-cent steps (in-tune notes read 0).

let loud = await a.stat('loudness')                       // LUFS
let [db, clips] = await a.stat(['db', 'clipping'])        // multiple at once
let spec = await a.stat('spectrum', { bins: 128 })        // frequency bins
let [min, max] = await a.stat(['min', 'max'], { bins: 800 }) // peak envelope for canvas rendering
await a.stat('rms', { channel: 0 })                       // left only → number
await a.stat('rms', { channel: [0, 1] })                  // per-channel → [n, n]
let gaps = await a.stat('silence', { threshold: -40 })    // [{at, duration}, ...]
let bpm = await a.stat('bpm')                             // 120.5
let beats = await a.stat('beats')                         // Float64Array [0, 0.5, 1, ...]
let { bpm, confidence, beats, onsets } = await a.detect() // full pipeline, one pass
let notes = await a.stat('notes')                         // [{time, duration, freq, midi, note: 'A4', clarity}]
let chords = await a.stat('chords')                       // [{time, duration, label: 'Am', confidence}]
let k = await a.stat('key')                               // {label: 'C', mode: 'major', confidence}
let coughs = await a.similar({ at: 12.3, duration: 0.4 })  // [{ at, duration, score }]: the other coughs

Meta

Property                         Description                                                                                                                        
.meta tags: {title, artist, album, year, bpm, key, comment, pictures, raw, ...}. Writable. meta.raw holds format-specific blocks untouched (WAV bext/iXML, ID3v2 frames, FLAC blocks).
.meta.pictures cover art [{mime, type, description, data, url}]. .url is a lazy Blob URL (browser) or data URL (Node).
.markers [{time, label}] in output seconds; edits shift or drop them.
.mark(time, label?) a marker at time, seconds of the audio as edited so far; {at, duration}, a region. Later edits carry it; in silence, at its distance from the sound nearest it; past the end, the end until a later edit makes the audio reach it. Chainable.
.regions [{at, duration, label}]; edits shift or drop them.

Parsed on decode, written on save; round-trips WAV, MP3, FLAC.

let a = await audio('song.mp3')
a.meta.title                     // 'Track Name'
a.meta.artist = 'Me'             // mutate
img.src = a.meta.pictures[0].url // lazy Blob URL

a.crop({ at: 10, duration: 30 })
a.markers                         // re-projected — outside markers dropped, inside shifted

await a.save('edited.mp3')        // tags + pictures preserved
await a.save('stripped.wav', { meta: false })   // opt out

Utility

Method                         Description                                                                                                                        
.on(event, fn), .off(event?, fn?) subscribe / unsubscribe.
.undo(n?) returns the undone edit, for redo via .run().
.run(...edits) apply ['type', opts] edits: op params (value, freq, …) plus range keys.
.dispose() release resources. Supports using.
Event                         Description                                                                                                                        
'data' pages decoded/pushed. Payload: { delta, offset, sampleRate, channels }.
'change' any edit or undo.
'metadata' stream header decoded. Payload: { sampleRate, channels, estDuration }: seconds it lasts, from the header where it says (WAV, AIFF, FLAC, an MP3's Xing or VBRI frame, a constant bitrate), else from its size; null when neither is known.
'timeupdate' playback position, as heard (~50 times a second). Payload: currentTime.
'play' playback started or resumed.
'pause' playback paused.
'volumechange' volume or muted changed.
'ended' playback ended: at its end, by stop(), or taken over by play({ from }); not in a loop.
'progress' during save/encode. Payload: { offset, total } in seconds.
a.on('data', ({ delta }) => draw(delta))  // decode progress
a.on('timeupdate', t => ui.update(t))     // playback position

a.run(
  ['gain', { value: -3, at: 10, duration: 5 }],
  ['crop', { at: 1, duration: 2 }],
  ['fade', { in: 1, curve: 'exp' }],
  ['insert', { source: ref, at: 2 }],
)
a.undo()                                  // undo last edit
b.run(...a.edits)                         // replay onto another file
JSON.stringify(a); audio(json)            // serialize / restore

Plugins

Method                         Description                                                                                                                        
audio.use(...plugins) register an @audio contract factory, a stat { stat, compute }, a codec { codec, test?, decode?, encode? }, a function receiving audio, or a registry name. Registry plugins need no use: a.compressor() and a.stat('truepeak') load them on first use.
audio.op(name, descriptor) register an op: a process function or { params, process, plan, resolve }.
audio.op(name?) one descriptor, or all ops.
audio.stat(name, descriptor) register a stat: (chs, ctx) => [...] or { block, reduce, query }.
.plugin(ref, opts?) a native plugin as an edit: VST3, CLAP, Audio Unit or LV2, by file, name or id, through the optional @audio/host (Node). Its parameters by key in its own units, functions or { t, v } to automate them; its latency taken off, its tail rendered (tail seconds, false none); notes, midi, bpm, timeSignature for what it hears; key a sidechain; preset, state; plugin one of a file holding several; params for a parameter named like an option (mix, at); isolate in a process of its own. audio --plugins lists them, audio --plugins NAME one's parameters. A Web Audio Module (WAM 2.0) by its module's URL or its class, in a page too: hosted by the WAM SDK in an OfflineAudioContext (Node: web-audio-api), the input rendered through it before playing, its compensation delay taken off, its decay until silence; parameters by id or label, automation, notes, midi and bpm as WAM events.
≡ Pedalboard load_plugin, a DAW's insert, a WAM host

Plugins also run without the engine: audio/batch over a whole signal, audio/stream over live chunks. Plugin tutorial.

import { compressor } from '@audio/dynamics-compressor/audio'
audio.use(compressor)                       // bring-your-own factory

a.freeverb({ room: 0.8 })                   // registry plugin: loads on first render; tail composes
music.ducker({ key: voice })                // sidechain via the key option
await a.stat('truepeak')                    // stat plugins land on a.stat()

audio.op('crush', { params: ['bits'], process: (input, output, ctx) => {
  let steps = 2 ** (ctx.bits ?? 8)
  for (let c = 0; c < input.length; c++)
    for (let i = 0; i < input[c].length; i++)
      output[c][i] = Math.round(input[c][i] * steps) / steps
}})
a.crush(4)                                  // custom op, chainable like built-ins

a.plugin('RX 12 De-click', { sensitivity: 6 })                  // native plugins: npm i @audio/host
a.plugin('AUDelay', { delayTime: 0.25, feedback: t => 20 * t })  // automated, on the timeline
audio(4).plugin('Surge XT', { notes: [{ time: 0, duration: 1, note: 'C4' }], bpm: 96 })

Worker

Call                         Description                                                                                                                        
audioWorker(source, opts?) same API, engine in a Worker; the main thread keeps a few-KB facade.
audio(source, { worker: true }) same, once audio/worker is imported.
{ worker: new Worker(url) } your own worker entry: codecs, plugins, your own code and messages, then audio/worker, which talks on a port of its own.
expose(a) → id, audioWorker.adopt(id, { worker }) hand an instance your worker made to the page as a facade.
audioWorker.context the page's AudioContext, the same as audio.context.

Across the boundary clip(), split(), clone() return promises; op errors emit 'error'; functions don't cross, use {t, v} curves. play() renders in the worker straight into the page's AudioWorklet: the main thread can stall for seconds without a dropout. Architecture.

import audioWorker from 'audio/worker'
let a = audioWorker('track.mp3')            // decode/edits/stats/encode in a Worker
a.gain(-3).fade(0.5)
let [mins, maxs] = await a.stat(['min','max'], { bins: 640 })  // transferred, zero-copy
a.play()                                    // rendered in the worker, played by an AudioWorklet (Node: @audio/speaker)

// your own worker: its messages stay its own, and what it makes plays on the page
import audio from 'audio'                   // worker.js
import { expose } from 'audio/worker'
self.onmessage = ({ data }) => self.postMessage({ out: expose(audio(data.file).gain(-3)) })

let worker = new Worker('./worker.js', { type: 'module' })   // page
worker.onmessage = ({ data }) => audioWorker.adopt(data.out, { worker }).play()

CLI

npm i -g audio, or without installing: npx audio …

audio [source] [transforms...] [sink] [options]

A pipeline: a source produces audio, transforms reshape it, a sink consumes it. The default sink is stat — printing an overview.

# sources
FILE         path, URL, or glob  ('*.wav' for batch)
-            stdin (or omit when piping)
record       capture from microphone

# transforms (chained left-to-right)
gain         fade        trim        normalize   crop
clip         remove      reverse     repeat      pad
speed        stretch     pitch       insert      mix
crossfade    remix       pan         split       resample
highpass     lowpass     eq          lowshelf    highshelf
notch        bandpass    allpass     vocals      dither
crossfeed    shrink      crossover   match       spectral
repair       copy        cut         paste       plugin

# sinks (terminate the chain — at most one)
stat [NAMES...]    print analysis (default)
play [loop]        open player UI
save PATH          encode and write (or `-` for stdout); `192k` bitrate, `24bit` depth

# options
-f --force         overwrite existing output
--format FMT       override output format
--macro FILE       apply edits from JSON
--cue FILE         split at cue-sheet tracks (with split)
--verbose          show progress
--help, -h         help (or per-op: `audio gain --help`)
--mcp              serve the CLI to AI agents as an MCP tool (stdio)
--plugins [NAME]   the native plugins @audio/host finds, or one's parameters

# a native plugin (npm i @audio/host): by name, its parameters as name:value
audio voice.wav plugin "RX 12 De-click" sensitivity:6 save clean.wav

# named options, after an op or sink: name:value
normalize -27 lufs ceiling:-2     ducker key:voice.wav     save out.m4a codec:alac

# compatibility shortcuts
-p ⇔ play     -l ⇔ play loop     -o PATH ⇔ save PATH

Playback

Audiojs demo

␣ pause · ←/→ seek ±10s · ⇧←/⇧→ seek ±60s · ↑/↓ volume · l loop · s save as · q quit

# play full song
audio song.mp3 play

# play fragment
audio song.mp3 10s..15s play

# play and loop a hook
audio song.mp3 30s..45s play loop

# play with effects applied live (streamable ops)
audio song.mp3 normalize broadcast highpass 80hz play

Edit

# clean up
audio raw-take.wav trim -30db normalize podcast fade 0.3s -0.5s save clean.wav

# scope a range (applies to whole chain)
audio in.wav 1s..10s gain -3db save out.wav

# range on a single op
audio in.wav gain -3db 1s..10s save out.wav

# filter chain
audio in.mp3 highpass 80hz lowshelf 200hz -3db save out.wav

# concat
audio intro.mp3 + content.wav + outro.mp3 trim normalize fade 0.5s -2s save ep.mp3

# crossfade into next
audio track1.mp3 crossfade track2.mp3 2s save mixed.wav

# voiceover
audio bg.mp3 gain -12db mix narration.wav 2s save mixed.wav

# music bed ducked under the voice (sidechain)
audio bed.mp3 ducker key:voice.wav mix voice.wav save episode.wav

# loudness to any target, true peak held; delivery settings
audio book.wav normalize -20 lufs save book.mp3 192k

# fix a video's sound, keep the picture
audio talk.mp4 highpass 80hz 4 normalize podcast save talk.clean.mp4

# master to a reference track (tone, width, loudness)
audio mix.wav master reference.wav save master.wav

# shorten pauses; the same cuts as an EDL for the video editor (.fcpxml, .otio too)
audio talk.mp4 shrink 0.3 save talk.edl

# split
audio audiobook.mp3 split 30m 60m save 'chapter-{i}.mp3'
audio album.wav split --cue album.cue save '{i} - {title}.mp3'   # cue-sheet tracks, tagged

# record
audio record 30s save voice.wav

Analysis

# overview (default sink)
audio speech.wav

# range overview — `audio FILE 0..10s` ⇔ `audio FILE stat 0..10s`
audio speech.wav 0..10s

# specific stats
audio speech.wav stat loudness rms

# tempo / beat grid / onsets
audio track.mp3 stat bpm
audio track.mp3 stat beats onsets

# loudness to spec: integrated, max momentary / short-term, speech only, true peak
audio mix.wav stat loudness momentary shortterm dialog truepeak

# pass or fail against a delivery spec (exit 1 on a fail); --json for scripts
audio episode.wav normalize podcast check podcast
audio chapter.wav check acx --json

# pitch / chords / key
audio song.mp3 stat notes
audio song.mp3 stat chords
audio song.mp3 stat key

# spectrum / cepstrum with bin count
audio speech.wav stat spectrum 128
audio speech.wav stat cepstrum 13

# stat after transforms (transforms apply, then stat)
audio speech.wav gain -3db stat db

Batch

audio '*.wav' trim normalize podcast save '{name}.clean.{ext}'
audio '*.wav' gain -3db save '{name}.out.{ext}'
audio 'chapters/*.mp3' check acx                      # a whole audiobook: one line per chapter

Check in CI

A file that misses its spec fails the build: one line per file, ✓ or the rule it broke, exit 1 on any fail.

# .github/workflows/audio.yml
on: [push, pull_request]
jobs:
  audio:
    runs-on: ubuntu-latest
    steps:
      - uses: actions/checkout@v5
      - uses: audiojs/audio@master   # or any CI: npx -y audio 'episodes/*.mp3' check podcast
        with:
          files: episodes/*.mp3
          spec: podcast              # streaming, broadcast, netflix, acx

Stdin/stdout

cat in.wav | audio gain -3db save -      > out.wav
curl -s https://ex.com/speech.mp3 | audio normalize save clean.wav

Tab completion

eval "$(audio --completions zsh)"       # add to ~/.zshrc
eval "$(audio --completions bash)"      # add to ~/.bashrc
audio --completions fish | source       # fish

FAQ

What formats are supported?
Decode: WAV, MP3, FLAC, OGG Vorbis, Opus, AAC, AIFF, CAF, WebM, AMR, WMA, QOA via decode. Encode: WAV, MP3, FLAC, Opus, OGG, AIFF via encode. Codecs are WASM-based, lazy-loaded on first use.
Does it need ffmpeg or native addons?
No ffmpeg: codecs and processing are JS and WASM. Playing and recording in Node go through a small native addon, prebuilt for macOS, Linux (x64, arm64) and Windows (x64), so nothing compiles at install; elsewhere playback falls back to ffplay, SoX or aplay. For the CLI, install globally: npm i -g audio.
How big is it?
In a page: 107 KB gzipped, the whole library minified (dist/audio.min.js); codecs and plugins load on first use via import(), so unused formats aren't fetched. In Node: npm i audio installs 14 MB in 294 packages (every codec and plugin), nothing compiled.
How does it handle large files?
Audio is stored in fixed-size pages. In the browser, with { storage: 'persistent' } (or 'auto', where OPFS exists), cold pages evict to OPFS when memory exceeds budget — auto-sized from navigator.storage.estimate() (quota/4, 64MB..512MB), overridable via {budget}; each instance keeps its own store, its copies share it. Works for decoded files, audio.from(pcm) and pushed streams. Stats stay resident (~7 MB for 2h stereo).
Are edits destructive?
No. a.gain(-3).trim() pushes entries to an edit list — source pages aren't touched. Edits replay on read() / save() / for await.
Can I use it in the browser?
Yes, same API. See Browser for bundle options and import maps.
Does it need the full file before I can work with it?
No. Playback, edits, and structural ops (crop, repeat, pad, insert, etc.) all stream incrementally during decode — output begins before the file finishes loading. The edit plan recompiles as data arrives, tracking a safe output boundary per op. Only ops that depend on total length (open-end reverse, negative at) wait for full decode.
TypeScript?
Yes, ships with audio.d.ts.
Does it have feature parity with FFmpeg / SoX / librosa?
Yes — the audiojs ecosystem covers the practical baseline of FFmpeg filters, SoX effects, librosa analysis, Pedalboard and MIREX, all as @audio/* plugins audio wires through one API (the few uncovered items are esoteric or deliberately skipped). Every effect, filter, generator and analyzer lives in the registry — call one by name and it loads on first use. Coverage matrix: docs/comparison.md.
How is this different from SoX / FFmpeg / Audacity / librosa / Web Audio / Tone.js?
In one line: audio is the only one that runs the same API in Node and the browser, with non-destructive lazy edits that stream during decode. The native tools (SoX, FFmpeg) are faster on raw throughput but have no JS API, browser, or undo; the browser libs (Web Audio, Tone.js, Howler) are real-time graphs, not file editors. Full feature and performance matrices vs pydub, librosa, aubio, essentia, Pedalboard, SoX, FFmpeg, Audacity and MATLAB are in docs/comparison.md.

Built with

MIT · ॐ

Projets similaires

Javascript audio library for the modern web.

JavaScriptaudioaudio-libraryhowler
Ggoldfire
25,4 k étoiles2,3 k

Headless Web Audio API

JavaScript
Aaudiojs
950 étoiles71

Audio waveform player

TypeScriptaudiojavascriptmusic
Kkatspaugh
10,4 k étoiles1,8 k