When Speed Outruns Judgement

How Creative Alignment Emerges

My earliest experiments with AI music made the speed impossible to ignore. Ideas that once took days to mock up could be heard in minutes. A collaborator could react to an actual piece of music while the conversation was still happening, rather than trying to imagine what a description might sound like.

At first, that felt like the obvious breakthrough. In film and television, time is always running out. Any technology that shortens the distance between an idea and something people can hear has real value.

The more I worked with it though, the less convinced I became that speed was the most important innovation. Generated music is capable of providing a convincing answer before the people requesting it have agreed on what the music should accomplish. A polished result can make that disagreement harder to recognize.

I had already lived through a version of that problem, long before generative AI entered the workflow. My composing partner and I were hired on a new television series after two individual composers had already been hired and fired. From the outside, we assumed the composers had been wrong for the show. Once inside the production, another possibility became hard to ignore: nobody seemed to know how to communicate about music, although everyone had very specific ideas about what it should do.

Our first major challenge was a bank-heist sequence full of tension, reversals, and shifting points of view. We took careful notes at the spotting session, interpreted the conversation, and wrote the cue. The written notes that followed came close to telling us which buttons to press at every moment: bring in a pulse here, remove an element there, make this moment bigger, change the sound at a particular cut. They rarely told us whose experience should guide the scene, what the audience should feel, or which dramatic idea the music was failing to express.

Some directions contradicted the spotting session and others contradicted one another. A request for a “big sound” could mean greater scale, more volume, a darker texture, a clearer turn in the story, or simply that the scene did not yet feel important enough. Each interpretation would lead to different choices in harmony, orchestration, rhythm, dynamics, and timing.

The showrunner had almost no vocabulary for music. He knew when something worked but could not reliably explain why, so he eventually turned music communication over to the picture editor. The editor knew the cut and had strong technical instincts, yet the notes increasingly prescribed sonic events without revealing the dramatic need behind them. We were receiving more information and understanding less.

After a couple of weeks of chasing our tails, we asked for a call with the producers, editor, showrunner, and network executives. The production changed the communication chain, and the showrunner took the reins again. His notes remained vague and sometimes almost comically imprecise. He might type the sound of a rhythm he imagined—as if “ratta-tat-tatta-tat” could tell us what he was hearing in his head—or use a phrase with no stable musical definition. The difference was that his reactions came from the person responsible for what the show was trying to be, and we had room to interpret them.

Over time, we learned his language. We learned what made him uneasy, where his attention went, and whether “bigger” meant more force or more consequence. His vocabulary never became especially musical; our understanding of the person behind it became more precise. The series lasted two seasons, and by the end all of us were proud of the work.

The weeks of contradictory notes delayed the creative discovery we needed. That discovery began when the production clarified who held the dramatic authority and restored interpretation to the composers. The goal was never a perfectly worded instruction. We needed enough shared understanding to make a purposeful musical choice.

Alignment means giving people enough shared context for their differences to become productive.

AI systems can collect everyone’s input, summarize it neatly, and generate music without establishing that shared understanding. The result may sound coherent because the system has combined the instructions successfully, even when the collaborators are asking the music to do different things. Speed becomes dangerous when it carries an unresolved disagreement all the way into a convincing cue.

Creative alignment requires context around each response: who offered it, what responsibility that person holds, which moment prompted it, and whether it describes a dramatic concern or proposes a musical solution. “The robbery should begin to feel inevitable” and “add a pulse at 01:14” are different kinds of input. A pulse may serve that dramatic intention, though it is only one possible answer.

That danger has changed how I think about CueMap, a scene-oriented system I began building to improve creative alignment in film and television production. In its current form, one user can gather the material relevant to a scene—plot points, dialogue, character motivations, timing, mood, instrumentation, tone—and generate a music brief written for people and an exploration prompt adapted to an AI music platform.

That risk applies to CueMap as well. If a collaborative version simply combines every response into one clean brief, it could hide the very differences it is supposed to help the team resolve.

A collaborative version would begin with independent responses from the showrunner, editor, producer, composer, and other members of the creative team. CueMap would preserve who contributed each idea, separate dramatic intentions from suggested musical execution, and organize the responses before combining them. The group could then see where people agree, where different language may describe the same need, and where the music is being asked to serve genuinely different purposes.

Some differences would require a decision. CueMap could return those questions to the person responsible for that part of the scene rather than quietly resolving them through synthesis. Once the relevant choices had been made, the system could produce a shared brief and exploration prompt. The resulting cues would let the collaborators hear the consequences of their choices and reveal whether further differences remained.

I have not built that multi-user version. Writing this essay has clarified what it would need to do. A useful system would reduce the time spent untangling avoidable confusion and preserve the time needed for interpretation and discovery. It would help the collaborators reach a decision before generated music makes one sound as if it has already been reached.

I think about how different those first weeks on the series might have been with a tool like that. It could have surfaced the contradictions between the spotting session and the later notes, distinguished dramatic intention from proposed execution, and shown that “big” was being used to mean several different things. We still would have needed the call, the showrunner’s judgment, the editor’s knowledge of the cut, and our experience as composers. We might also have spent those weeks discovering what the show needed music to do instead of figuring out how to talk to one another.

None of that shows up in the brief or the prompt. It shows up in whether our choices allow the music to express what the scene couldn’t say on its own.

Previous
Previous

Before Anyone Else Hears It — Part Two

Next
Next

Where Does the Creativity Live?