A Million Little Pieces Of My Mind

The History Of Sound Recording

Fake Stereo!

By: Paul S Cilwa Posted: 9/2/2026 Page Views: 163
Hashtags: #Science #History #Music #SoundRecording #Stereo #ArtificialIntelligence
How we finally solved the one-microphone problem.
Estimated reading time: 24 minute(s) (5635 words)

I've spent a good part of the last few decades building a digital music library of my own. Streaming services can drop an artist overnight—a licensing dispute, a contract lapse, a public relations problem—and the album you've played four hundred times simply isn't there any more. So I've ripped thousands of CDs, a wall of LPs, and even a trunkful of 78s and 45s. And what came off those older (pre-1958) discs is mono. Which got me wondering what it would take to hear those recordings the way the people in the room heard them.

Real Talk for the Squad

(That's Gen-Z for "A Message For Persons Born Since 2000.")

Since I just used two pieces of jargon that stopped being common knowledge somewhere around the Reagan administration: a 78 is a shellac phonograph record that spins at 78 revolutions per minute and holds about three minutes of music per side. Three minutes. That's it. That's the entire reason the pop song is three minutes long; it's not an aesthetic principle, it's a manufacturing limit that everybody got used to and nobody ever revisited.

Shellac deserves a word of its own, being the least likely substance ever asked to carry a symphony. It's a resin secreted by the lac insect, a scale bug that swarms certain trees in India and Thailand and encrusts their twigs with the stuff. The twigs are scraped and the resin washed and dried into amber flakes. You have almost certainly eaten some: it's the glaze on shiny candy and the coating on a good many pills, and it's what your grandmother's dining table was finished with.

A 78 was never pure shellac, though. A typical pressing was mostly filler— pulverized limestone and slate, carbon black for color, cotton flock for strength—with the shellac merely binding the aggregate together. And the grit wasn't an economy measure. Steel needles were sacrificial, meant to wear down to match the groove within a play or two, and the abrasive in the record is what ground them to shape. You threw the needle away and put in a fresh one. The record did the sharpening.

All of which imposes a ceiling that has nothing to do with how well the performance was captured. A stylus dragged across compressed rock flour makes noise, and it makes it worst in exactly the register where the detail lives: cymbals, brushes, string overtones, the s at the front of a word. Every 78 has a hiss-and-crackle floor some thirty decibels below the loudest thing on it, which means the quietest sound a 78 can hold isn't set by the microphone. It's set by the gravel. And the same abrasive that sharpens the needle is chewing the groove while it does it, taking the high frequencies first—so a well-loved record is a duller record.

Those YouTube videos in which 1920s records are played on modern turntables and sound, well, not awful—the records they are playing, were not beloved by their original owners; if they had been, the grooves would have worn to almost nothing within a few years of purchase.

Quieter materials existed but kept losing in the marketplace. Edison pressed his Diamond Discs on a phenolic surface over a wood-flour core, played with a permanent diamond stylus. It lost anyway, and not on sound. Edison cut his grooves vertically, so his discs wouldn't play on a Victrola and Victor's wouldn't play on his machine—and Victor had Caruso. Edison personally auditioned the artists, and his taste ran narrow and conservative—little use for jazz, few stars signed, and he marketed the machine rather than the performer, at exactly the moment Victor was inventing the celebrity recording artist with their Red Seal series.

Vinyl became available in the thirties and got its real trial during the war, when the lac came from Asia and the shipping didn't: the V-Discs shipped to the troops were pressed in vinyl because there was no shellac to be had. It still took until 1948, and the introduction of the "LP", the Long-Playing record, for the industry to give up the gravel.

But those V-Discs, with their quieter surface noise, have given us wonderful examples of what a 78 could have sounded like if the materials had been better. Sadly, they returned to shellac after the war—the vinyl had been a shipping decision, not an audio one, and shellac was cheaper.

A 45 is the vinyl record that replaced it in 1949—seven inches across, spinning at 45 RPM, with a hole in the middle big enough to push a banana through, with one song per side. The good song was the A side. The B side was whatever the label had lying around, though, occasionally, the B-side would surprise everyone and become the bigger hit.

The big hole in a 45 wasn't a design flourish. It was sized for the fat spindle of a record changer that could stack a half-dozen singles and drop them one at a time—the 1950s equivalent of a playlist.

Both formats are monaural ("mono" for short): one channel, one signal, one loudspeaker's worth of information, from mon, the Greek for "one," and aural, meaning "of or relating to hearing." Stereo didn't reach the consumer until 1958, and it took most of the '60s to become the default. Which means that essentially everything recorded before Eisenhower's second term exists only as mono—including most of the music I actually want to listen to.

Not that stereo was a new idea in 1958. In 1881—four years after Edison's tinfoil phonograph—Clément Ader lined the stage of the Paris Opéra with telephone transmitters and ran the wires across town to the Exposition of Electricity, where listeners held one receiver to each ear and could tell where on the stage a singer was standing. His Théâtrophone went on to run as a paid subscription service until 1932. The idea then sat unused for fifty years, because knowing how to transmit two channels is not the same as knowing how to record two channels and keep them in step.

Except that I'm not sure the gap was as empty as that. I grew up in St. Augustine, which had a museum called the Oldest Store—a turn-of-the-century general store preserved with its merchandise still on the shelves. Among the exhibits was a cylinder phonograph with two reproducers, each running by its own rubber hose to its own earpiece, playing a cylinder cut with two separate bands of grooves. In the sixties it was labeled as a stereo machine.

We weren't allowed to play it, so I can't tell you what it sounded like. But consider what recording meant before microphones: a horn, and sound loud enough to drive a cutting stylus on air pressure alone. Two bands of grooves means two styli, which means two horns, which means two positions in the room. Whatever was on that cylinder, the two bands cannot have been identical. Stereo wouldn't have been an achievement. It would have been unavoidable.

The Bee In My Bonnet

(That's Boomer for "The Thing Living Rent-Free in My Head.")

Recently I've been writing and producing an album called 1944, set on an alternate Earth where I was not only alive in the forties, but leading a big band on a USO tour. Getting the arrangements right meant getting the sound right, and the sound that came back was a revelation. Listen to "Wee Willie Winkie" and you're hearing a room: a band arrayed across a bandstand, brass to one side, reeds to the other, a singer standing out in front of all of it, and the hall itself wrapped around the whole arrangement. I can't swear it's exactly what a 1944 audience would have heard. But it's believable in a way no surviving 1944 recording is, because the surviving recordings were made through a single microphone onto a single track.

Which is not a new impulse. Starting in 1970, Time-Life sold a mail-order series called The Swing Era that attacked this exact problem in the only way then available: they hired studio orchestras, transcribed the original arrangements off the original 78s note for note, and had them played again in stereo. Billy May did much of the conducting, while listening to the original recordings through headphones as he did so; and a good many of the players had been in those bands the first time around. If you wanted to hear the Swing Era in stereo, somebody had to go back into a room and perform it again.

The sets sold enormously (including to me), and collectors have been ambivalent about them ever since. The playing is clean, the stereo is real, the charts are correct—and something is missing that nobody has ever managed to name precisely. It may be that the recordings are too perfect. A transcription captures every note a man played and nothing of why he played it that way; and the men reading those charts in 1970 were thirty years older than the ones who cut the originals, and a good deal more careful.

In any case, that made me greedy. I've got Glenn Miller and the Andrews Sisters and Harry James in my library, recorded on only one channel, and I would very much like to hear them the way the people at the Hollywood Palladium heard them. Which turns out to be a problem the recording industry has been failing to solve, in escalating ways, for about seventy years…until now.

What Two Ears Are Actually For

Before we can talk about faking stereo, we should be clear about what stereo is trying to fake. And the honest answer is: not much. You have exactly two ears, they're about seven inches apart, and they are pressure sensors. Neither one knows where anything is.

What your brain does with them is compare. There are three clues, and only three:

  • Timing. A sound off to your left reaches your left ear before your right ear. The maximum difference—a sound directly off one side—is about 0.7 milliseconds. Seven ten-thousandths of a second. Your brain resolves differences down to about ten microseconds, which is a genuinely absurd degree of precision for a piece of wet tissue.
  • Loudness. The far ear is a little quieter, because it's farther away and because your head is in between.
  • Tone color. This is the subtle one. Your skull blocks high frequencies much more effectively than low ones, so the far ear hears a duller version of the same sound. And the folds of your outer ear filter sound differently depending on whether it arrived from above, below, in front, or behind—which is how you can tell a plane overhead from a truck behind you, using ears that are both pointed sideways.

Everything else—the sense of a room, of distance, of a stage with depth—comes from reverberation. The direct sound arrives first; then the reflections off the walls, floor and ceiling arrive over the following tens and hundreds of milliseconds, and the pattern of those reflections tells you the size and shape and hardness of the space you're standing in. You do this constantly and without noticing, and you're good enough at it to hear the difference between a bathroom and a stairwell with your eyes shut.

What Stereo Actually Is

Two channels. That's the whole definition. Two independent signals, kept separate all the way from the microphones to your ears, so that the differences between them can do the same work your ears were already doing.

The trick that makes it useful is the phantom center. Put identical signals in both speakers and you don't hear two sources; you hear one source, floating in the air midway between them, where there is no loudspeaker at all. Make the left one slightly louder and the phantom slides left. Delay the right one slightly and it slides left again. Which means that with two speakers and a pair of knobs you can place a singer anywhere along a line between them—and if you get the reverberation right, at any apparent distance behind that line as well.

Headphones, Loudspeakers, and Surround Sound—Oh, My!

Everything above assumes loudspeakers, and loudspeakers give you something headphones can't: crosstalk. Sound from the left speaker reaches your right ear too, a fraction of a millisecond later and slightly muffled, exactly the way real sound in a real room would. Your brain gets the arrival-time and tone-color cues it evolved to use, and the illusion sits out in front of you where a stage belongs.

Headphones deliver the left channel to the left ear and nothing else, ever. No crosstalk, no head shadow, no room. The result is that the stereo image collapses into a line running through the middle of your skull, which is why hard-panned sixties mixes sound so unpleasant on headphones, while merely quaint on speakers.

The exception is binaural recording, made with microphones in the ears of a dummy head. It sounds startlingly real on headphones—and oddly flat on speakers, because the speakers then add a second set of head-shadow cues on top of the ones already baked into the recording.

Once you understand that stereo is a two-channel trick played on a two-sensor system, the surround formats stop being mysterious. A 5.1 system has five full-range channels—left, center, right, and two behind you—plus a low-frequency channel, which is the ".1" and which carries only the deep rumble your ears can't localize anyway. Dolby Atmos adds height, and stops thinking in channels at all: the mixer places an object somewhere in the room and the playback system works out which of your speakers should produce it.

But none of that gives you more ears. Five speakers don't produce five-channel hearing; they produce five sources, and your same two ears still have to sort out the combined arrival times, levels and tone colors. More speakers means fewer phantoms and fewer compromises—a sound directly behind you is very hard to fake with two speakers in front, and trivial with one (or more) speakers behind. The mechanism never changes. All any of it does is give your existing hardware better raw material.

Phasers On Stun

A sound wave is a cycle of rising and falling pressure. Phase is simply where in that cycle a wave happens to be at a given instant, and the phase relationship between two copies of the same wave is the difference between them.

Delay is phase. If you take one signal, copy it, and delay the copy by half a cycle, the copy is 180 degrees out of phase—its peaks land on the original's troughs. Add the two together and they cancel out to silence. Delay it by a quarter cycle and you get a partial cancellation. This is why phase is not an abstraction: it's the mechanism by which a delay turns into a change in what you actually hear.

Now notice what that means. The primary clue your brain uses to locate a sound—the arrival-time difference between your ears—is a phase relationship. Which is enormously tempting, and which is where the fakery begins. If a few hundred microseconds of delay is what makes a trumpet sound like it's on your left, then surely you can take a mono trumpet, delay a copy of it, and put the trumpet on the left.

You can. Sort of. What you can't do is put the trumpet on the left without doing the exact same thing to the clarinet, the piano, the drums and the singer, all of which are stuck in the same mono signal. And because a fixed delay cancels some frequencies while reinforcing others, what you get is a comb filter—a series of notches carved through the spectrum, which sounds like the whole band is playing through a length of drainpipe. Worse, the moment anybody sums your fake stereo back to mono for a car radio or a television broadcast, those notches become permanent.

Stereo Before Anybody Knew What To Do With It

Two-channel stereo reached the record-buying public in 1958, and I've written about how the discs themselves worked in Listen To The Music. The engineering was elegant. The mixing, for about seven years, was terrible.

The problem was that stereo arrived as a sales feature before it arrived as a craft. Customers who had just paid for a second speaker wanted proof that it worked, and the surest proof was to put something in one speaker and nothing in the other. So early stereo mixes hard-panned: rhythm section on the left, horns on the right, singer wherever there was room. Early Beatles albums are notorious for it—an entire band on one side and a lead vocal on the other, with a hole in the middle you could drive a Volkswagen microbus through.

The purest specimen I own is the Smothers Brothers' Two Sides Of The Smothers Brothers from 1962. One side of the album is a live club performance; the other was cut in a studio, and on that studio side the entire orchestra is in the left channel and both brothers are in the right. Not "weighted toward." In. Tom and Dick are standing in one speaker and the band is playing in another speaker several feet away, and the two rooms have nothing whatever to do with each other.

What's missing isn't the panning. Real orchestras really do put the brass on one side. What's missing is that in a real room, every instrument reaches both of your ears, but not in the same millisecond, and the reverberation of the hall is shared by everything in it. Cross-channel bleed and shared ambience are what glue a stereo image together. Those early mixes had neither, because each instrument had been recorded on its own track in its own isolated booth and then assigned to a channel like a seat on an aircraft.

"Electronically Reprocessed For Stereo"

Meanwhile the labels had a catalog problem. Every record made before 1958 was mono but stereo was what sold, and reissuing a mono record into a stereo market felt like leaving money on the table. (On the other hand, the number of consumers who actually had two speakers in 1962 was still small enough that the labels didn't dare release only stereo versions—so they usually released both monaural and stereo versions of the same record, and the stereo version was often a different mix entirely, and sometimes even a different recording.) So they started manufacturing stereo out of recordings that never had any, and printing a phrase on the jacket—"Electronically Reprocessed for Stereo," "Enhanced for Stereo," Capitol's trademarked "Duophonic"—that collectors learned to read as a warning label.

The techniques arrived in roughly this order, each one an attempt to patch the failure of the last.

1. Delay And Phase Shift

The first idea was the obvious one: split the mono signal in two, delay one copy by ten or fifteen milliseconds, and sometimes flip its polarity for good measure. This does widen the sound. It widens it the way a funhouse mirror widens your mother-in-law. Everything smears, the comb filtering gives the whole record a hollow, phasey quality, and if the polarity was flipped the record partly disappears when played in mono—which, in 1962, was how most people were still listening, and another reason why monaural records continued to be made, and sold.

2. Split The Spectrum

The second idea was to divide the frequency range instead of the time base: roll the treble off the left channel and the bass off the right, so the two speakers carry measurably different signals. Capitol's Duophonic did this, usually with a short delay layered on top.

It fails for the same reason all of these approaches fail: Nothing has been placed anywhere. The bass fiddle isn't on the left; the low half of the entire band is on the left, and the high half of the entire band is on the right. The singer's chest tone is coming from one speaker and her consonants from the other. It doesn't sound like a room in which musicians are playing. It sounds like the sounds of musicians have been shattered into pieces and glued back together, wrong.

3. Add Reverb

The third idea was smarter, and it's the one that survives in modern plug-ins. Feed the mono signal into an artificial reverb—a plate, a spring, a chamber, later a digital algorithm—and generate two different reflection patterns, one for each channel. Now the two channels genuinely differ in ways your brain recognizes as spatial, because uncorrelated reverb is exactly what a real room produces.

And it does work, up to a point. The result sounds spacious, which is a real improvement over sounding like a row of slots. What it isn't, is located. The direct sound—the part that carries every localization cue you have—is still identical in both channels, so every instrument still sits in the phantom center. You've built a hall and put the entire band on a single point in the middle of it. And reverb smears transients, so you buy your spaciousness by softening every drum hit and every consonant on the record.

By the early seventies most labels had quietly given up. Reissues went back to being labeled "Mono" or, more often, "Original monaural recording"—which by then had become a selling point rather than an apology, mirrored a few decades later by "audiophiles" who happily paid extra for inferior vinyl copies of popular recordings for their perceived "warmth" and "presence," which were actually the result of the same distortion that had been driving the early stereo experiments.

What The Curtain Told Us

Starting in the 1920s, researchers began running comparisons in which listeners heard either live musicians or a recording, with an acoustically transparent curtain hung across the stage so nobody could see which was which. The early results were humiliating for the musicians. Audiences routinely preferred the narrow, rolled-off, filtered sound of a phonograph record to the actual musicians standing eight feet away. Edison had been building marketing campaigns on the same principle since 1915—his touring "Tone Tests" pitted a live singer against one of those Diamond Disc phonographs I described earlier, and the company insisted that audiences couldn't reliably tell them apart.

The conclusion drawn at the time was habituation. People didn't prefer accuracy; they preferred what they were used to, and by 1930 what they were used to was records, either played on a phonograph or over the radio.

Harry Olson at RCA went after that conclusion in 1947 and complicated it. He put a live orchestra behind the curtain and a mechanical acoustic filter in front of it, so that the restriction was applied to living musicians rather than to a recording. Under those conditions the preference flipped: listeners wanted the full frequency range. Olson concluded that the earlier results hadn't measured taste at all. They'd measured distortion—the wide-range equipment of the day was so nonlinear that extending its range mostly extended its ugliness, and listeners were reasonably voting against that.

He was probably right. And it still doesn't dispose of the point, because the habituation effect kept right on showing up every time the technology changed. Audiophiles preferred the warmth of vinyl to early CDs. Producers spent the nineties adding tape saturation to digital recordings that had none. A generation raised on 128-kilobit MP3s learned to like the sizzle, and a generation after that mixes records to sound good on a phone speaker. What you grew up with is not what sounds accurate. It's what sounds right, which is a different measurement entirely, and one no engineer has ever been able to argue anybody out of.

The Argument

Which brings us to the fight, and it is a real fight, conducted with more heat than the stakes strictly justify.

On one side are the preservationists. A 1938 recording is a historical document. It was made with particular microphones in a particular room by engineers making deliberate choices, and the mono mix isn't a deficiency in the document; it is the document. Every intervention is a layer of somebody else's opinion between you and the artifact—and the entire history of "improving" old recordings, from Duophonic through the noise-reduction era, is a history of confident people destroying things simply because they were old. Once a reissue has been rechanneled, de-noised and re-equalized, the label frequently throws out the original it started from.

On the other side are the reconstructionists, and their argument is that the document was never the point. Nobody in 1938 sat down to make a mono recording. They sat down to make a record of an event that was in a room, in three dimensions, and the mono was a limitation they'd have abandoned in a heartbeat if the technology had existed. What you're preserving when you preserve the mono mix is not the performance. It's the equipment.

My own view is that this stops being a fight the moment you stop insisting on one answer. Nobody thinks a colorized print of Casablanca should replace the black-and-white negative. Plenty of people are glad both exist. Keep the flat transfer—that's the document, and it should never be overwritten, deleted, or "upgraded." It should be preserved, in the archaeological sense: cataloged, stabilized, and left alone. Then make whatever else you want, and label it. The sin isn't the reconstruction. The sin is passing the reconstruction off as the original.

Stemming The Tide

Which brings us, finally, to the thing that changed: Artificial Intelligence.

By now you may have some vague idea of how A.I. works on text. Text models work on tokens: a finite vocabulary of word fragments in a sequence, where the model's job is to predict what comes next. Audio has no vocabulary. A CD-quality stereo recording is 88,200 numbers per second, none of which mean anything individually, and all of the structure you care about—pitch, rhythm, timbre, who's playing what—is smeared across thousands of them at once.

So music models generally don't work on the waveform directly. They work on a spectrogram: a chart of which frequencies are present, at what strength, at each moment in time. Turn a recording into a spectrogram and you've turned a wall of numbers into something with visible structure—a bass note is a stack of horizontal lines at the bottom, a cymbal is a vertical smear across the top, a voice is a set of harmonics that bend and wobble together. Which is a problem shaped very much like image recognition, and image recognition is the thing machine learning learned to do first.

A stem is one component of a mix, isolated: just the vocal, just the drums, just the bass. In a modern studio they exist by construction, because every part was recorded on its own track and the finished record is those tracks added together.

What they don't exist as, is a thing you can extract from a finished recording. Once tracks have been summed, the sum is all there is; there's no more a "vocal channel" hiding in a mono file than there's a red channel hiding in a bucket of purple paint. Which is why source separation was, for decades, a research embarrassment. Every filtering approach fails on the same rock: the singer and the trumpet occupy the same frequencies at the same time, and no filter can pass one without passing the other.

And yet we humans can pick a voice out of a trumpet solo without the slightest effort, in exactly the mixture where every filter fails. Psychologists call the problem auditory scene analysis; the familiar version of it is the cocktail party effect, where a roomful of people all talking at once arrives at your eardrums as one pressure wave and yet you can follow a single conversation out of it with little or no effort.

The rules your brain uses have nothing to do with frequency bands. Harmonics that are all multiples of the same fundamental note get heard as one voice. Sounds that begin at the same instant get assigned to the same source. Partials that wobble together stay together—a singer's vibrato bends every one of her harmonics at once, and that shared motion is a signature no trumpet in the same octave is going to match. Add a lifetime of having heard trumpets, and you aren't separating the mixture at all. You're imagining what must have gone into it.

The machine-learning approach doesn't filter. It recognizes. Train a model on tens of thousands of multitrack recordings, where both the finished mix and the individual stems that built it are available, and the model learns what a snare drum looks like on a spectrogram, and what a bowed cello looks like, and—critically—what a snare drum looks like when there's a cello playing at the same time. Given a new mixture, it predicts what each source's spectrogram must have been. Tools like Spleeter and Demucs will hand you four or five stems from a finished record in less time than it takes to play it once.

The distinction matters, and it's the source of both the power and the problems. The output is not a "filtered copy" of the input. It's a reconstruction—the model's best account of what was probably there. When the model is well-trained it's uncannily right. When it isn't, you get artifacts: a watery, underwater quality where it wasn't sure, or cross-bleed where the mix was too dense to untangle. And on a 1944 recording—made through one microphone, band and singer and hall all committed to one track at once, with 78 RPM surface noise sitting on top of everything—the A.I. will be a good deal less sure than it is on a modern pop mix.

For this week, anyway.

Building A Stage Out Of Stems

But once you have stems, the seventy-year-old problem simply evaporates.

Every one of those failed analog tricks failed for the same reason: the machine had no idea what was in the signal. A delay line can't put the trumpet on the left, because a delay line doesn't know a trumpet from a bass drum. It can only do the same thing to everything.

Separate the recording first, and you're no longer processing a mixture. You're processing parts. Which means you can do to each part precisely what a real room does to a real instrument: give it its own arrival-time difference between the ears, its own level difference, its own head-shadow filtering appropriate to the angle you've placed it at. Then run all of it through one shared reverb—a convolution reverb, built from an actual measured impulse response of an actual hall, so every instrument is in the same room as every other instrument. Then add the parts back together.

The result is stereo in the only sense that matters. The trumpet isn't louder on the left; the trumpet is on the left, with a full set of the physical cues your brain has been using since infancy, and the hall is wrapped around the whole band the way a hall is supposed to be. It is doing exactly what your ears do, in reverse.

It is also, and there's no point pretending otherwise, an invention. The model guessed at the stems, and somebody—or something—chose where to put them. Glenn Miller had a seating chart, and my software doesn't know what it was. What comes out is a plausible 1944 bandstand, not "the" 1944 bandstand, and any separation artifact in the stems gets baked permanently into the result. Feed it a clean 1958 mono master and it's remarkable. Feed it a worn 78 and it can produce something that sounds less like a big band than like a big band being described to you by someone who was standing outside.

For this week, anyway.

Which is what I would do if I were the Library of Congress. Keep both. The mono master is the historical record, and it should be preserved.

However, I am more a listener, and I'd rather listen to stereo, good, believable-sounding stereo, than to the historical record.

The purists are right that the mono master is the historical record and that nobody should be allowed to overwrite it. The reconstructionists are right that a document of a performance was never the same thing as the performance. Both of those are true at once, and they've only been arguing because the technology forced a choice between them. It doesn't any more. Storage is free, and there's no longer any reason a recording can't be both a preserved artifact and a living performance, as long as you keep the labels straight.

Meanwhile, now that I've reminded myself of it, it's time to remaster that Smothers Brothers album's studio recordings, and put them back in front of the orchestra where they belong.