SC-063 · Expert
The Future of Music Creation: AI's Role in the Industry #aiinmusic
Guest: Dr. Kobi Abayomi, Statistician
Summary
Dr. Kobi Abayomi, a statistician and former Warner Music Group data science leader, argues that generative AI models are trained on copyrighted music and that creators need attribution and payment. He explains how models compress music into a latent space and how prompting can produce market ready songs.
The legal picture is unresolved. The US Copyright Office says AI generated artwork is not copyrightable, while Tennessee's ELVIS Act restricts voice cloning. Abayomi calls attribution a solvable problem, using metadata and blockchain identifiers so outputs can carry royalty signals back to rights holders.
The discussion highlights that generative AI compresses the time and money needed to reach professional quality, which threatens small and midsize careers. Abayomi does not see the shift as a total loss: human connection, storytelling and consumer choice may matter more, and the pressure could force modernization of copyright and revenue channels.
As of the episode's release on 9 July 2024.
Key takeaways
- 01Generative music models are typically trained on copyrighted recordings because public domain music sounds different from what Western listeners expect.
- 02The US Copyright Office says AI generated artwork itself is not copyrightable, while Tennessee's ELVIS Act restricts voice reproduction without permission.
- 03Attribution is solvable: metadata and blockchain identifiers can trace outputs back to training data, so monetized works can return royalties to rights holders.
- 04Host Jakob Wredstrøm's prompting tests produced one in five usable generations, showing AI can already reach market quality after enough iteration.
- 05Small and midsize music careers face the most risk because AI output can match early-career producer quality, but human connection may remain the differentiator.
Chapters
- Preview
- Welcome and Kobi's background
- Generative AI's industry moment
- Why training data is contentious
- How US copyright treats AI art
- Career sustainability for creators
- Prompting and latent space
- Xerox analogy and attribution
- What will matter
- Outlook for creators
Guest
- Dr. Kobi Abayomi, Statistician
Questions this episode answers
How do generative AI music models learn from existing songs?
They ingest large corpora of sound and metadata, often including copyrighted music. In the gold standard approach, people label songs by genre and features, then the model learns a latent space from that data. Models cannot produce beyond what they have been exposed to.
Is AI-generated music protected by copyright?
Currently no. The US Copyright Office has determined that AI generated artwork itself is not deserving of copyright, and it remains open ended whether human authored work with prompting is copyrightable. Tennessee's ELVIS Act separately restricts voice reproductions without permission.
Can artists get royalties when AI models train on their music?
Yes, Dr. Abayomi argues attribution is straightforward. He envisions adding blockchain identifiers to training data so every generated output carries metadata that can be scored, and then payments flow back to the original rights holders when the work is monetized.
Will AI replace human musicians and producers?
Abayomi does not think generative AI is a sum minus for creators. It adds another layer of complexity to an already difficult digital transformation, but it may accelerate the need to modernize copyright law and revenue channels, which could be virtuous.
What happens next for AI and music copyright?
Abayomi expects competitive legal and pseudo legal actions, including agreements between rights holders and AI firms. He thinks pressure from AI could accelerate modernizing copyright and revenue channels, which may be virtuous for creators.
you can't construct this space without the use of the material in the first place
Episode notes
Generative AI is transforming music creation, revolutionizing production, creativity, and copyright. In this episode, Dr. Kobi Abayomi explores its implications for the music industry and its potential to shape the future for creators, offering key insights into the intersection of technology and artistry.
Highlights:
- Examining AI's influence on music creation and industry changes.
- Understanding the intersection of technology and artistry.
- Addressing the challenges and opportunities AI presents to musicians and creators.
Topics
- Generative AI in Music
- Music Copyright
- Training Data
- Creator Livelihoods
- Blockchain Attribution
- Latent Space
Transcript
Transcribed automatically. Names and terms may be misspelled. Every line is timestamped: select a time to play from there.
Read the full transcript
And I think of an analogy here, you know, Xerox machine, say you, you know, take some piece of art and you copy it and you, you know, rub your hand over it or something like that. And then you take another one, which is sort of similar, you rub your hand over that too, and then you tell the Xerox machine to sharpen the two of them. You can't claim that these two things are the progenitors or parents of this child object that you created. That's just dishonest, if you ask me. The training data part of it is incredibly important. These models don't produce things beyond, you know, what they've been exposed to.
Me and Dr. Kobe go very deep into generative AI. What does it mean to train generative AI? What does it mean to have an output? What do they look at? How does it work? And what does it mean for industry? Will people get paid? Well, we talk about all of this in this episode. Thank you for listening. Hey guys, and welcome back to Sound Connections Podcast. Today, we are joined by Dr. Kobe at two at night from California. Welcome, doctor. Hey, how are you doing?
I don't know how to use those formal titles. We don't have them much in Scandinavia, but I've been looking very forward to this. We're talking about something that we've covered on the podcast before. It's AI, but we're talking about something very specific. It is about data input from generative AI, how to think about it, how to understand the concepts, what are the consequences, how does the process work? And, Kobe, you are an expert in this field. To set the scene, could you let people know who you are and what you do?
Sure. You know, part of it, it can be a way of introduction. Usually, when people say, you know, Dr. Kobe or Dr. Amayomi, I go DJ Leroy to my friends. That's a little joke. But, actually, I was a DJ in undergrad and graduate school. I had a radio show, 89.3 WBAI, Barnard College Radio. DJ Leroy, love to the world experience. I went to grad school for statistics. I'm a statistician.
And I was a professor for some years and then left to go make my way in the corporate world. That part of my life has been mainly around, I would say, advertising, technology, but really on the supply side. I worked for some B2B companies and then media. So, I spent some years working on data science for digital media, found my way to Warner Music Group where I created and led for some years a data science function, which brought me to music and, you know, I love and was amazing.
And so, I continued in that. I work now for a startup that is basically an AI company for people who have rights or are making money from music rights to optimize listening patterns. And then, along the way, you know, picked up some tricks in machine learning relevant to the problem at hand, which is understanding how people listen to stuff and how to squeeze the most use out of it. Yeah, and it's a complex space, AI, because there's so many nuances to it and there's so many different types of way of talking about AI.
And one of the things that obviously is the center of attention right now in the music industry is generative AI. You know, it has been in the public view for some time, especially within, you know, graphics, arts, that kind of stuff it's had its way. And we see a lot of moves happening with, you know, video and Sora. And now, especially with Suneo and Udio coming out very publicly and making this accessible, there's a lot of concerns about, you know, the training data. That's always been a concern with these companies, but now it's affecting the music industry more than it used to be.
And before it was a, you know, a problem for another day. And now it's a problem for now. So, but the basics of sort of addressing the issues is understanding the issues. And that's sort of one of the things I want to linger on a bit today. But let me understand a bit before we go into it. The story you work for now, what is the nature of that AI work? What is the approach to AI there? Now, I would uncover the generative AI company. This is more about predicting listing patterns from music as information.
Music has a, I mean, everything does, right? It has representation as information. I should say everything. There may be some things which are unrepresentable in the sort of numeric or information system that we have, but let's just say almost everything does. And music has a very nice one. People have thought about this for a long time. The signal processing literature, for example, is quite old, right? And you can go all the way back to people like Claude Shannon, right?
Or even Boltzmann, you know, as far as statistical physics and people's ability to understand communication processes, et cetera. Sound and this relationship with statistics, at least from the 1960s, right? Like where the National Academy of Sciences had gathered a bunch of literature, a lot of people out of Carnegie Mellon, honestly, to work on problems that speech detection and things like that. And so there's been a long history of intersection between computer science, mathematics, statistics, and sound.
And you find these things so interesting. You know, the way the world seems to work is, you know, things pass through similar sort of modes, but at higher levels, right? And so now we're passing through this mode where we're, you know, readdressing sound in particular. Obviously, art and more generally, but now with more sort of a computational ability to, and I put in quotations, create things that are compelling, you know, songs, information gathered, prediction output from models that sounds like a song or looks like a picture, right?
And models that are conversant enough or low code enough that accept regular input. And so it's now we've come to the point where we're addressing sort of ordinary people and their ability to sort of intersect with creativity. And, you know, a lot of things are springing up around that, but this is really just, you know, the latest iteration of a long process between, uh, information and our ability to sort of reproduce it. It makes sense. And one of the, one of the, the big issues with where we are now is, uh, recently if people have followed, uh, Suno just raised $125 million, um, in a series B, I believe.
Um, and the thing about that is the music industry is looking at that, so, okay, there, there will most likely be a significant economy surrounding generative AI and, you know, most likely without, you know, probably any doubt that it's trained in something that's not licensed. So there's someone here not being paid, but also combined with the conflict, conflicting nature of what is the future of the music industry right now. That, that makes it super sensitive. Could you, before we go deep into it, could you, could you break down what are the problems with generative AI when it comes from, but from a copyright standpoint?
You know, so let me, after that, let me just sort of say a couple of things just in, in general about technology and, uh, music, right? They've had a long sort of career together, each advancing and, and sort of, you know, there being a sort of rear guard fight against it the whole way through, um, from the invention of the photograph right where people were like, you know, how are you going to have this device that's going to fit in people's houses that's going to make people solipsistic and introverted. Um, you know, some of these, you know, some of these, they had more or less sort of readily apparent versions of, uh, new ways of, of making music that are familiar to the people who are making music at the time.
And some of them seem revolutionary. So you've always had this tension between, uh, sort of the virtue, if you want to put it like that, of technology and the response of the music community. Um, and along with this tension, you've always had this, I would say always, but let's just say in the recent history, um, this notion of sort of speculation of people who were sort of crowding into the field to sort of make their money, you know, back in the, uh, the advent of, of the recording industry would be the people who were turning sheet music and sort of folk music, black music into, um, recorded objects.
Right. And let's get this all down, the tin pan alley and stuff like that. And now it's, you know, um, speculators, VCs, companies like Sona. Um, so I spent some time reading, uh, some stuff from the copyright office and patent office. And then, so patent doesn't apply here. Right. And it's, it's not the, it's not the area of, uh, American Jewish prudence that addresses, um, music and music reproduction.
But it was one of the things I found in reading the latest stuff from the patent office on AI is they have a much more, how should I put it? Um, I have the word that's come to mind is loving, but it's the word I want to use, but, um, the, uh, a much more sort of accepting or, well, let's just stay with loving take on, on the discipline AI itself. That they did a study last year on how many patents they're putting out and have AI in it and how can they sort of support the field and have it grow.
And then I moved over to the copyright side, which, um, you know, these are both government, you know, organizations, um, corporations and, um, the copyright side is much more adversarial, uh, and, and the, and the, and the writings on that, um, where the U S is and just taking the U S just because, you know, that's where I live and that's sort of where I know best is right now that, and let's put the train data to the side for a second is that an AI generated artwork itself is not deserving of, of copyright.
Uh, and, and they've left a little bit open-ended, um, whether or not human authored as a human authored can, can mean from, you know, I wrote the program myself. Or I use a program, uh, you know, that detects prompting, you know, like that to create something. And that's sort of, it's a little bit open-ended I say, um, as to whether or not human authored work will be copyrightable at this moment. The answer is no, uh, we've seen some others, I would say governmental action around this last year.
Tennessee passed an act on, on name and likeness voice in particular, right? Uh, the Elvis act. I think the thing was titled, um, basically restricting like, Hey, um, there are no copyrights available for, uh, um, voice reproduction. So you can't just, you know, take like something else can do and train on a voice and then you go and now you're singing like Dolly Parton. So that's, you know, sort of reboting without express, uh, permission. And then, um, you wouldn't, one would imagine royalty transmission.
So some of this I think is just, uh, let me say it this way. I think some of this is, is how we intersect with just the notion of creation itself. And one of the themes I'm, I'm getting in, in reading about this stuff is there's a, it, this sort of idea of sort of determinism or sort of randomness.
And if there's a mechanical process, right, the, you're divorced from, from the body, if you will, that can go through some rules and do this and make this thing. It seems that, uh, the, the decision is to come down that these things aren't copyrightable, but things that appear to happen more randomly. One of the things that I, that I read in some of the copyright, the U S copyright office literature is that if something were created as if it were taught to a child and a child reproduced it.
Right. So there's this personalization of these AI models, um, and, and they're all, and, and the sort of, I won't say desire, but sort of the, at least the utility of using their opacity as in some way, mirroring sort of human learning processes of those things seem to be sort of more acceptable, or at least more artistic in a sense that, uh, deserving of, of copyright, um, period paragraph.
So that's sort of the, the apparatus developing around all of this. What do I think is going to happen and how it'll all, uh, be, you know, filtered out. Uh, these are going to be, I think, you know, competitive, um, actions in the legal, uh, or pseudo legal space. And by pseudo legal, I mean, you'll have, uh, agreements between, uh, rights holders and, uh, you know, generative AI firms that they either just come up with themselves or ones that they're forced to, uh, because the, you know, the case will either go in one way or another.
Um, and, and so that'll be a whole other thing that, that needs to sort of fall out of this. It's very interesting. It's very complex. And it's, um, I think this is sort of what confuses people the most is the nuances of the different interpretations, the different ways of thinking about it. And there is no clear cut presidents really, uh, at large across all cases that can, you, we can point to and say, this is probably where we're going. So there's a lot of discovery to be made right now. Um, that's interesting because, you know, ultimately that will speak something about, uh, let's call it the market share that, you know, generative music will take or add.
We, we don't really know how it will affect it, but most likely there's going to be a, this is my assumption of some minus for creators, uh, right now, as it looks, um, that especially, let's call it the small to midsize careers that quite quickly can be affected by the quality. Um, like I, I experimented a lot with generative aspects of music and images and videos. And recently I've been doing some, I've been building some AI avatar assets, like almost like a R and D kind of low level process.
And some of the generative music that I'm, that I'm reaching out now, which is basically just practicing, prompting and prompting and prompting the structures and what works is at the level that, you know, most producers wouldn't be able to hit within the first six, seven years of their career. Right. Which means there's going to be, um, this is what, this is my biggest issue. My biggest issue with this is the incentive of building a career is that there needs to be some element of sustainability to it. You cannot basically just slave of, you know, I'm going to build up for seven and eight years and hope at the end that it's going to be some income associated to it.
So, so the, the period of, I want to do this too, I can actually do this is quite essential. That's right. And there's needs to be some revenue associated to it. What happens with generative AI, AI right now makes that space even more difficult is my assumption. Right. Yeah. That's sort of the model and like performance in general, right? One would hope, uh, outside of other reality TV shows is that there's a period of sort of practice that you have to go through to differentiate yourself as an artist.
And then you're rewarded, you know, for that art because of your ability, your proficiency, your virtuosity, right? And this sort of seems to short circuit it. Having said that, and just sort of speaking to where the models are now, and you, and you mentioned it yourself, you're, you know, going through the process of learning stuff and it's not straightforward. You know, so there's, I was reading, um, and I, maybe I can play it for you later, I don't know. Uh, there's this, this blog that I read and, uh, they've been experimenting with, um, generative AI songs.
So they're trying to come up with songs that are actually listenable. I mean, one of the things you'll find anybody who's using these models is that they're using the word people use is hallucinatory. The process for generating, uh, new output has a bit of randomness in it. And that's sort of intentional. That's sort of the jittering that happens with these models. Um, in order to get, they, they got an album of, uh, folk music and dance music, which I do have to admit is, is listenable.
Um, you know, even good, the dance part, especially, which is, you know, perhaps you're more used to hearing that sort of stuff in dance music, but they said they had to run through 30,000 different iterations of songs, right. To get 15 that were good. So at this point there's, you know, so there, there's still, um, you know, some Rubicon to be crossed if you're going to, to do something that's going to be compelling and, and remunerative. Um, which is. No, to add to that, because, uh, I'm, I'm very obsessed about prompting.
Uh, that's one of the things that I spent a lot of time too, because there's some internal logic and I have, I have now, uh, it took a lot of time and iterations like them. I have now broken down in theory, what are the specific parts of this prompt and the structure of the text that works. And I can now, without exaggerating one out of five generations are now to the quality that this could be placed in the market, uh, which is scary. Like, uh, yesterday I was playing around with, with the new theory I have behind the prompts and I've tested it.
And in a matter of 20 minutes, I had four songs where like, I would have no problem with these songs. And that's, that's when it gets scary for me. Like it's, it's just, if it's a matter of prompting, like, and, and the technique that could be, you know, shared, educated, uh, and, and right now it's difficult. Like those 30,000 iterations I had over 1,500 myself. Um, but, but there is, there is a way, there's a structure and it's repeatable at least that. Yeah. That's, that's called steering.
Um, uh, so there's several papers out about the ability to steer language models and without optimization because you're addressing the model at the sort of, you know, I'll say steady state. It's not really, um, and the prop that you're doing, that you're doing, be able to figure out which part of the, uh, latent space you're addressing, um, is actually, I don't know if this is the way we would jump right on this is, um, uh, a precursor, if you will, to the deeper sort of study of attribution, right?
Like this piece of data was used either in this way or in this sort of quantum, uh, in this output and therefore needs to carry some sort of royalty recognition. And, uh, steering is one way and prop engineering is one way to try to uncover that, right? Yep. Yeah. And yeah, I've, I've been working a lot on that, but, but let's, let's go sort of to, to the basics because this is really interesting and there's so much to talk about, but I do want to understand the basics.
Let's just take a generative AI music company that people would know by now. What is, what is the assumption on what has happened with getting into that place when it comes to training data? Oh, they've trained it on everything they can find, uh, copyrighted, uh, and not, uh, you know, the thing is, especially if you're trying to, uh, market something to, you know, the Western year, right?
Um, the Western popular year, you've definitely trained it on copyrighted music, you know, the, the space of music that's not copyrighted or in sort of the public domain is, um, it sounds different, right? So if you do it on just that stuff, the stuff you're going to get out, it's going to sound like it and not going to be compelling to people who use it. So these things have all, I, you know, I say with very high confidence that these things have all been trained on copyrighted material.
Uh, yeah, that, I think that's the general consensus of the music industry, but it's very nice to hear from your mouth that's also your assumption. So, so what, what goes into that? What, what does it mean to train a model on material in general? Sure. It would depend upon the setting. It depends on what you're trying to get out of it, but more or less, uh, you can consider, uh, anything in the class of neural networks. Um, and I'm just wanting to speak about this because we're talking about these type of models, um, as a sort of input output, uh, and the output is something that's going to be similar to the sort of stuff you're going to want it to output sort of as it's used in perpetuity, whatever.
And, and the input can be a, any number of things. So let's take one of these, uh, um, you know, text to sound sort of models, right? So the input will be sound, um, metadata, which sort of encapsulates or as a notion of the thing that might be prompted upon later. And the output will be, um, the song involved. So some of this can happen automatically, you know, people are getting sophisticated enough to be able to process corpora, uh, under different sort of principles, um, sort of self training sometimes.
But the standard or, I wouldn't, I don't know, as they'd even say standard, but I'd say maybe gold standard way is to ingest something and then to have real people spend some time, uh, labeling it, right? So in the, in the, in, in the gold standard way, lots of different songs, um, this is a jazz song. This is a so-and-so song, like we're sophisticated enough with processing sound data that you could grab from another source sort of attribution on the sound.
Somebody has already decided the information for this sound sits here in jazz and has these features. Um, but to really, really do it, you know, with the chef's kiss perfection, you'd have somebody pass through that and add other stuff. It's like, you know, David Fanborn and, you know, IAuto, stuff like that. So the training data part of it is incredibly important. These models don't produce things they can't, uh, beyond, you know, what they've been exposed to.
And, well, one of the explanations that I've, I've had recently is sort of how they reproduce it and sort of how they copy. Uh, this might not be correct, but I'm just trying to paraphrase at least what I've understood is that, um, the AI, if that's, that's probably not the right term, but like the machine learning model or whatever you want to call it, looks at the, looks at the song and the interconnectivity between the data points rather than the actual data.
Uh, if, if that's somewhat correct, can you, can you come back with a better explanation than I just did? Sure. So let's, let me, by the way of doing this years, years ago, um, back when I was grad school, I've worked this, uh, I was a professor for one year at a small little arts college. And I had a really good student who wanted to study hidden Markov models and, um, hidden Markov model, say you're trying to predict stock price, right?
It's like price one day, next day, next day, next day in Markov model is a model which has behind stock price. There's something else going on, right? We don't see it or we don't know what it is, but if we estimate that really well, um, what we see in the stock price is quite predictable, you know, some levels of probability. And one of the ways in which we had to estimate the, uh, hidden Markov model was sort of pass through a time series once and estimate, uh, these states and then sort of pass backwards to tune them.
And that is analogous to a very simple, uh, neural network in, in the sense that there's a layer, there's, uh, information in the layer that's sort of latent, which is the word I'm trying to get to, and there's information that's sort of seen, which is the thing you're trying to predict. So what you are talking about is what's called the latent space, uh, in which the model operates.
Now, these models may have very many layers, um, a particular version of one of these types of models called an auto encoder, which has a series of layers sort of going in to a latent space. It's in the series of layers coming out, but more or less you're taking, uh, representations of things and projecting them onto a space where now the space is informative, right? Uh, I'm thinking of very old versions of this, like say you were doing this in linguistics and you had, you want to learn a space where the word king and queen aren't that far apart from each other.
So that there's sort of some semantic thing that you can grab out of this as meaningful. And that's what's happening is that the model is going back to the space that is learned from the data, right? Uh, and then in that space, grabbing stuff that matches the prop, the prop directs it to where in the space to go, uh, and then, you know, pull something out. Um, and that's sort of where the intelligence comes in, right? It's sort of looking at things, making conclusions, packing them together and having understanding based on all the things you've looked at.
What is, you know, the, one of the things that, and you know, when you talk to the people who spent their careers in neural networks is that they even express amazement. Yeah. It kind of works, right? That if I could come up with this space, which, Matt, you know, meets some condition that, you know, I can get something, uh, useful out of it. And so, um, you know, this, I wanted to bring up the hidden mark on model because I wanted to say something like this is, again, these, this is the next iteration of stuff that people have been doing for a very long time.
They have been very different, you know, sort of adjacent types of, of modeling. And this is the latest, most sophisticated iterate, you know, iteration of it, but not wholly different from sort of what's been done before. Um, yeah. So one of the arguments that I've heard is because these models look at the, again, I'm just going to butchering with my own words, the interconnectivity between data rather than data itself. It, it has a argument for not, um, needing to compensate, uh, the copyrighted material.
Um, uh, yeah, you know, I mean, somebody who's doesn't want to compensate copyrighted material will say something like that, but I said, to me, that sounds ridiculous. Like, um, you can't construct this space without the use of the material in the first place. Right. And in fact, the more sparse the space is, the weirder the stuff is, that's going to come out of it. Um, so the space needs to be dense and the space becomes dense either from, uh, you know, being dense to the training data or having really strong probabilistic assumptions, which again, may make things weird.
And I think of an analogy here, you know, Xerox machine, uh, you know, you say, you, you know, take some piece of art and you copy it and you, you know, run or rub your hand over it or something like that. And then you take another one, which is sort of, uh, similar, you rub your hand over that too, and then you tell the Xerox machine to sharpen the two of them. And you can't claim that these two things aren't the progenitors or parents of this child object that you created.
That's just dishonest. Yeah. If you ask me, um, yeah, I mean, another way to say that, this is why, you know, I, I spent so much time reading the patent literature, which is almost celebratory of the stuff is that there's art in creating the mechanisms and tools, which allow this to happen in the first place. Right. Um, and that's recognizable as well. Now, the part of it, that this is problematic, it was just, it's exactly what you were saying. I'm so glad you introduced, uh, your thoughts that way is we're for to recognize, you know, the art and sophistication of say the Xerox machine, then who do we recognize, right?
The creator of the Xerox machine, the, you know, the technicians who refine it, uh, or just the guy who comes up and flops down to copy. And so I think that's something that needs to be hashed out. Um, I, you know, part of that, I think will be people arguing back and forth and part of that, I think will be the way in which, you know, the public receive these works of art, right? There's a big difference between, I, years ago I used to have this, uh, Wurlitzer, uh, organ and you know, it could play songs, you know, pass the button, play me a Roomba beat and we'll play a little Roomba beat and then we'll go, you know, go about, but big difference between that and somebody getting down who's an accomplice organist and can really play.
Um, maybe these tools will get sophisticated enough so that they're close to this organism who can really play without any, um, art being put in it. But yeah, but that's, that's where it's going to file down to. Yeah. And, uh, oh, I was sitting yesterday and, uh, I was home with the kids and I was, I had a prompt idea. Uh, and actually, you know, it's, you know, I, I cracked the sort of prompting, uh, you know, a few weeks ago, but then I had just like this idea of extra refinement and I went back to the computer and I prompted it and it worked and I tried again and it worked and I tried again and it worked and like, okay, I have done 30 generations.
I have 10 successful songs. Uh, and then I went back to, okay, what if I use the same theory on the lyrics? Like I had some, you know, things that I've written myself and then sort of use that. And then I did some prompting and prompt engineering or steering, as you called it, and then put that in like, okay, it works. What, what maybe yesterday was a very overwhelming day for me because, um, I'm almost, I'm a bit disheartened. Uh, not so much for the industry.
Like we'll, we'll find a way to, to make something relatively healthy out of the situation. But the whole base of what we do is art. And right now, what I see happening is that the, the product itself could be, I don't want to say replaced, but it could have, it could be very clear comparison between someone who has put in blood, sweat and tears into something.
And that's something that's pretty generated and being equal of nature for the, for the ordinary consumer. Um, that speaks to, okay, maybe there's some other elements we need to double down, down on, but like there's so many creatives in this space who have sacrificed so much, way too much to be where they are. Now to see that livelihood will be challenged so severely, um, and that sort of, it's not really anything to do directly about the talk today, but it does have a lot because the compensation that maybe could be in place could maybe alleviate that pain a bit.
If not a lot, then be a part of the, the package of alleviating the pain of what's to come. Right. Well, let me say a couple of things. The first of this to sort of help alleviate pain and it could just be lies. I'm telling myself, um, well, the first part, I don't think it's the second part, maybe the first part is attribution. That's not a difficult problem. And I think it's a red herring that is treated as such. And it actually makes me think of some of the motivation of the speculation behind it, that it's sort of proposed as a problem.
We can't figure this out. Oh my God, if only we could figure out, you know, whose name is on this box, then we could deliver it. That's not a difficult problem. There's so many ways to incorporate rights attribution in the data or to detect it, uh, from the output side. There's, there's several papers, um, that are, there are looking at the semantics of the latent space, which would direct you back to the, uh, you know, the data, which generated that part of it.
There's a fixing metadata. You know, we spent years thinking about NFTs and blockchain, but somehow we've forgotten that now that we've gotten to attributing data across a, a derivative AI supply chain. Uh, so that's the, that's a red herring. That problem will be solved in as much as people desire or to be solved. But the second part of it, you know, a lot of what you're saying, and I'm not saying that this is going to be exactly the same sort of regime shift is similar to again, what people said back when, when we're moving from sheet music to photographs or from acoustic instruments, electronic ones.
Or even recently when we moved from sort of point of sale to this, you know, digital sort of smorgasbord. And the truth is that you, we have seen shifts in the way in which people consume art, some of it not virtuous. I think we're still dealing with, um, the digitalization of music and how to figure out how to monetize that better, uh, for the artists. And it definitely has had effect on the art that's been created. The, you know, the sounds songs are getting shorter and shorter and shorter to, to match that mode, uh, the way in which artists site contracts, a single, more single base, you know, to match that sort of mode.
So the, you know, the field itself, the, the way in which artists are able to, you know, get paid or rewarded definitely changes, um, sometimes in virtuous ways, sometimes in desultory ways. But all that is to say that I don't think it necessarily has to be, uh, all bad. And then for, and for these reasons, at least one, I think engaging people in the craft of music. Now this is stipulating that we're going to have a regime where we're going to be attributing works fairly and training things fairly.
Right. Basically this is an appreciation of history. Right. Um, I, I think that's a good virtuous thing. Right. And there are ways to get people more into music as they gamify it or play with things. Um, these track libraries, these digital audio work systems. I was looking at this company, Bria, I don't want to give anybody a plug, but, you know, allows you to sort of change your voice. Um, those things are attractive and interesting. And from what I understand, they're trying to do it in a virtuous way by paying the voice actors, you know, up front who they're training also.
And the second thing is, I hope, I would hope that all of that generates a deeper appreciation and need for music and everyday life. Um, there are so many other ways in which we experience or could be experiencing music. Right. Um, one of the things that's troubled me about the digitalization of music is a lot of it has become reduced to that, you know, live music, being around other people, experiencing music in community, having these sounds sort of envelope us, you know, at festivals.
Um, there's another company that's sort of like an Airbnb for music that sort of sprung up and what they do is have these micro venues and have artists go to micro venues and you stay on this sort of trail. Like I heard this on Tuesday and Wednesday at somebody's house. So I think there's ways that we can take advantage of, you know, a deeper sort of investigation of music that this, that this stuff, these sort of facile tools may provide to sort of feedback into a deeper sort of appreciation of music and musical experiences that can't be, you know, replicated by.
Yeah. And also I believe there's, um, there's a, there's a different, there's always been different roles in creativity, uh, and to make it super like super simple to understand. I, I believe that there's the, there's the storytelling facilitation aspect and then there's the execution aspect. Right. And sometimes they're in a creative, they might be in, in one person, sometimes they're in separate people. And one of the things, again, I, I have a master's in music production, but I, I was never a great musician.
So my, my biggest struggle was that I was always dependent on other creatives to get my work done, which means there was less money to me at the end. Uh, and that just didn't turn out to be a very viable business model. So I went into doing stupid stuff, like entrepreneurship. Um, but, but what I'm seeing now is my facilitator mentality and my executive producer mentality is so applicable to AI because I've learned to understand how to tell the stories. And I've learned how to communicate to partners, how to act in certain ways, in certain contexts to get the optimal out of them.
It's basically the AI is my musician. You know, my guitarist is Chad Chibiti and my bassist is Claude and my, my mixing engineer is something different. So the, the way of thinking about system processes, contributors towards one narrative is what's happening. So, um, that is what I see happening is people who have that skillset will be the creators that will thrive in this environment.
And people who do not have that skillset are the ones who are going to struggle. Um, I, that, that I can't, I, you know, honestly do not disagree with. I think that this is a real opportunity for all of these pieces to enhance music production. I love the way you just said it because you know, you, you're like a composer and a storyteller. Like I need to get this to sound like that. I need that to sound like that. So you're already used to sort of putting together these musical elements, right. And constructing a piece from it, um, I, you know, that's, I try to be, you know, Bob Marley has a, uh, a thing.
You have no fear for atomic energy, right. Because none of them can stop the time. And I always find that, you know, I think that I believe that if people have the right motivation and sentiment that they'll find their way to virtuosity, even, you know, as tech, I don't think technology is bad, right. Um, I think it's a good thing for producers to be able to have more tools. I, you know, people were up in arms. Oh my God, auto-tune, right. Um, the earlier version or parts of the versions of these, uh, at least sound based AI models rest on the machine learning models that finally got better at being able to extract, you know, stems from songs, right.
Which was a difficult thing to do, just sit down in front of an analog track board, like, oh God, I got to get rid of that and push that. So that, so all those things are useful. They may, you know, and again, I wanted to jump on, I hope that deepens the desire for music and musicality societally, right. And the biggest questions that comes from, from that talk is, okay, so what is going to matter? Uh, and going back to sort of the story I had about yesterday and prompting this music is I can objectively say that, uh, I, I do, we, we own a few different companies in my venture studio that sort of expand a lot of this space.
And if one of the given productions that I made yesterday was made by producer, I would sign them immediately. And that's the level now, like the, the only thing, honestly, the only thing that lacks doubt is audio quality. Like it, when, when you perfect the prompt, it's audio quality and you know, I have a hard time believing they're not going to address that. So, so the question is, so what is going to matter? Uh, and that's sort of what I'm left alone with because when or how, and will it happen?
Absolutely. It will happen. And it's going to be much more sophisticated than what we believe. So what is going to matter? Right. And I believe at the end of the day, it's a human connectivity is what's going to matter. And then we need to figure out what results in that. How do you create a human experience? And that I believe is what we're left with to a big degree. You know, I, I don't, I don't want to, one and one should never either leave their wallet out in front of, or play some bet in the belief of, you know, the, the virtuosity of the music business or industry rule number 4080.
But, uh, you know, this is one of the things that I've heard from people who spent a lot of time in the music industry. And so the, uh, the, the part of it that relationship centered and, um, divorced from the production, that's sort of savoir faire. Um, but so maybe the, some of that has a role to play here, uh, in, in the way in which this next sort of regime is crafted.
I'm also reminded of this quote by this, um, I'm looking at the quote right now. Our author and video game enthusiast, Joanna, message you ska. He says, I want AI to do my laundry and dishes so that I can do art and writing, not for AI to do my art and writing so that I could do my laundry and dishes. Right. Um, and I'm also reminded of, uh, in the eighties, there was a television show, Max Headroom. It turns out that Max Headroom wasn't, you know, as strongly CGI generated as it appeared, but Max Headroom was this CGI S character is really an actor who played it, um, inside a TV set.
And I think it might've been like, it's a detective show, like sort of campy and they would go on adventures. He'd be stuck in the, in the TV. And I remember when that came up, people were like, oh my God, all the TV shows are going to be like this with Max Headroom and CGI generated. And, and as it turns out, no, um, that there was a limit to the amount of consumption, uh, that people were willing to devote to that mode of art. Uh, and then lastly, I completely agree with you. Yeah, you're right. The sound quality and the artifacts, I think, uh, to spend a lot of time, uh, some of that, honestly, I, some of that, and I've thought about this before, um, back when I was at WMG is the training corpora that are being used.
Right. I think, um, as these models become pruned and able to be retrained more easily on more narrow corpora, that's one way of addressing the sound quality and the artifacts you hear. Uh, so that's something I think we'll see in the next iteration. I think you, you're touching an interesting point, uh, and now I'm projecting back to what I think you said, is that the consumers in the CGI example were the ones choosing not to necessarily appreciate that, that form of expressing oneself.
And that's one of the arguments I hear a lot, and that is, uh, there will be an active choice from the consumer to participate in AI in a specific way and, uh, elements that has a stronger human presence. The big question there is how for the consumers to navigate that, how would they know, how would they not know, and if that's through trademarks or whatever, like that's, what I'm hearing again and again is the belief that humans will take a stand standpoint and an active choice of what to consume and not to consume.
Not based on perceived quality, but beliefs, but because of the origin of the material, um, and what needs to be put in place is a way for consumers to navigate that space. That's a good point. I can think of examples where that information, you know, affected the, uh, the consumption of something. I'm thinking of Milli Vanilli, honestly, right? And, uh, once it became, they became exposed, oh, you know, verboten, even though people loved it before that, that same thing happened now.
Uh, if you were to lay, this is a, you know, completely AI generated, you know, when people have reservations about it, I don't know. Um, but I also think that, um, that discussion isn't strictly supply, you know, side. And then audience on both sides and just for the people who are creating as well, their sort of willingness to, uh, constrain what they offer or not, or have all the information included in what they offer or not.
Right. I, and I, I take this, and so that this, which is, can be virtuous, needn't always be. I'm thinking of a conversation I had with a record executive some years ago, uh, about the need for sort of forecasting and understanding listing patterns. Uh, and the response I got was like, oh, why do we need to know that? You know, it's all random. It's just what the kids like. And I, my response to that was like, you can't tell me that drill rap is some natural thing that just came from what the kids like, right?
It's just, if you constrain, if the people, you know, if you're only offering people dirty water, they're going to choose, you know, amongst the choices of dirty water. And the same thing here, if all we're going to offer people is going to be synthesized versions of art and all they have to choose from is that, um, you know, how do they discriminate? Yeah. It's interesting. Um, there's so much to uncover in this topic, uh, but obviously we don't have unlimited time.
Um, it's, I've really appreciated you to some degree breaking down the processes of how they train data. Um, but, but before we leave, I need to understand the output a bit as well. So we've talked about the training material, how it goes into this. I forgot the word you used something with L. It's very pretty. Uh, yes. Latent space. Latent space. Um, and sort of how it, uh, at least how I understood it is that the output gathers the information if in that latent space, space and producing output.
Can you, for my sake in the audience to say, explain the output a bit, a bit more short for sure? Well, let's, let's just go back to the Xerox example again, and, um, let's just refine it a little bit. Say, here you are from the Xerox machine. You, Xerox-esque machine. You've got a bunch of, um, pieces of art. Um, these art have been gathered because a prop came in, uh, through the vacuum tube. And said, I, uh, Monet or whatever, right?
And so you gathered a bunch of stuff. Um, a little bit of randomness in that. Some of the stuff is actually Monet. Some of the stuff is not. Copy all the stuff. Buh, buh. It goes into the buffer of this Xerox-esque machine. Now, inside the Xerox-esque machine, inside this buffer, there's a bit more, I hesitate to say transformation. That's a strong word for that. But I would say jittering or noise. Uh, generally noise is used when, for the normal distribution, Gaussian distribution, that's added to, uh, what's happening inside the machine.
And then it spits out something, which is a combination of all of the things that it gathered, uh, that were relevant to the prop that came to the vacuum tube. And this bit of, sort of, noise to, sort of, smooth things over and, sort of, you know, concatenate them all together. And that's more or less what's happening, uh, for output. Um, you're going to a place, the late space, that's been trained, updated in the first place.
Uh, and then pulling out a representation that's having a bit more, sort of, refinements, you know, probably too strong a word to use. Uh, I don't want to just say jittering or randomness at it, but somewhere between that, you're, you're adding some distributional information for the thing that you get out so that it's sort of, it's presentable. And that's more or less what's happening. So, you know, in other words, let's take it sort of from a musician. You have a musician that lives, listens to, let's say, Bob Mali.
He remembers that, he has an impression that it is. And then someone asks him, hey, can you play me some Bob Mali? And then he goes back, oh yeah, that was sort of like this. This is sort of how I remember it. And then he sort of performs it. So it's not actual Bob Mali, but it's my memory of what Bob Mali might be. And then I pull it out. Is that too simplified? Yeah, the only thing I want to say about that, because I've heard that explanation used to justify not paying Bob Mali rights. Like, you know, and the answer I would say to that is, if you're playing a Bob Mali song, you know, everyone knows you're going to have to cough up some money to Bob Mali's estate.
Yeah. Yeah, so it's still, because that, you know, that this is where we have the discussion, right, with these technology companies is they argue, you know, that memory we've created, which is personal memory or, you know, conclusions based on the input. We're not using the actual input. We're using our interpretation of the input to generate the output. Again, a red herring. I envision, and with some rapidity, right, this could happen tomorrow, where the data that go in to train the latent space are attributed with, say, blockchain attributes.
So everywhere in the space, you have some additional metadata vector that carries along with it these blockchain identifiers. When something is created from there, it's yanked out, the identifiers come with it, and then they're scored as to sort of the strength of their presence in it, right? And then whenever that thing is monetized, back to the blockchain, hit with the scoring vector, you know, money goes back to whoever the original right holder was. Is your assumption that all of these, all or most of these generative AI models are built on blockchain?
No, I'm saying that could be done. I'm saying that could be done, right? Like, so it's not, so what's, the start of it is, you know, where people are saying, well, you know, we're not pulling out that exact, you know, file. And so, therefore, I can't pay you, and I'm saying, you're not, okay, you're not pulling out the exact file, you're pulling out several files from which this file cabinet was constructed. Add the extra data, which includes the file names, if you will, blockchain, and when you pull out this object from this, you know, file cabinet, it'll have with it all the attribution that could be blockchain that will allow you to, yeah.
No, there's no, as far as I've seen, nobody's done this yet, but I'm saying it should be straightforward. Yeah, I know, I know several companies that I know personally that are, that are working in this space, and I very much agree to your approach of the technological solution. The, the big question is whether or not we can get access to sort of the previous, uh, training systems and how to integrate with them if there's a willingness to do so, and, yeah. Well, you, you know, if, if people are unwilling to be virtuous, I think at some point, this is going to be something that's going to be under the heading of discovery.
In a court case, right? That's interesting, okay, that gives me some confidence. Well, we've, uh, had an interesting talk, uh, you've been very, very generous, um, interestingly enough, this is a space that is just hard to understand, right? Because neural networks, AI, even though, you know, there's, there's fairly logical explanations of what it is and how it works, it is so new for so many people. And, you know, we barely got to know a blockchain when the NFT hype was there, and then we thought that was over.
So, like, I think we're still very stuck with sort of basic core understandings of how technology works in that space. And then when you add the level of complexity that this does, it, it confuses people a lot. But it's, it's very nice to hear that there might be the most possibly are technological solutions that through either virtuous opening of data or by caught will be open up for to accredit the rights holders. Um, and this is, this is for me a lot. Kobe, what's your last comment on this interview?
Um, sure. What's your outlook on this space? How will, do you believe it's a sum minus for the music that's been creators? Or do you believe it's a sum plus? I don't, I don't think it's a sum minus. I think that, uh, as far as creators, and I have some friends who are like, you know, real musicians, they've been in a, uh, you know, depression, um, you know, value, if you will.
Uh, for a while around this digital transformation, right? Before we even got to AI. I think it's been hard for people to navigate that and people are figuring that out. I think this adds another layer of complexity. I don't think it necessarily, you know, deepens the depression. Um, perhaps, you know, the epitaph of all this is that it will accelerate, uh, this, the tent of tent that I think is looming between artists and creators.
You know, that we've, we've perhaps seen precursors of, and, you know, artist strikes and things like that, uh, where people are going to have to elucidate how to modernize copyright laws. Right. And, uh, and revenue, uh, channels. Uh, and this may, this may quicken that, right? Whereas people sort of passive last 10 years of Spotify, but like now, okay, now it's time we really got to do this. Um, so that may be an outcome. It may be virtuous.
Um, I, you know, again, I don't, I'm not anti, I'm not the technology person. I, all of what you just said before that I thought was really beautiful. You're right. It's a lot to learn and it's, um, moving quickly. I, at even this age, the age of my career, I spent a lot of time reading new papers so that I can learn, so that I can, you know, operate. Um, it's an exciting time. Um, excitement always brings highs and lows. Well, it's been a pleasure.
I'd love to have you on. Uh, now it's a bit later, it's 3.15, so you can go, go to bed again or continue working. I'm going to go wake my cat up and tell him to move over. Amazing. Thank you, Kobe, for your time and we'll talk soon. Thank you, Jacob. Take care.



