[11:02.650 --> 11:05.610] Singing and the human voice. [11:09.350 --> 11:11.610] On to the pitch analyzers. [11:12.270 --> 11:22.630] The pitch analyzers are a command line utilities for frequency analysis focusing on four string violins and their pitch range from G3 to E7. [11:24.250 --> 11:33.710] We had this idea to analyze pitch using Python and the ParcelMouth library which provides a Python interface to the Pratt software's internal code. [11:34.230 --> 11:42.050] Pratt is a free and open-source software package that is commonly used in linguistics and phonetics to analyze and synthesize speech. [11:42.670 --> 11:45.570] The use case here is to review a student's tuning. [11:47.090 --> 11:48.330] So what this... [11:52.640 --> 12:01.220] ...draft a pitch analysis script using a generative AI prompt to design a Python 3 based command line application for this purpose. [12:02.360 --> 12:10.000] When it came to the code review session, Andrew pointed out that the results were a little bit too reliant on frequency analysis. [12:11.060 --> 12:18.260] We also got somewhat hesitant to use the ParcelMouth library further after realizing it is primarily used for voice analysis. [12:19.120 --> 12:26.520] The jitter and shimmer values we initially believe to help with our analysis are actually used to analyze human vocal chord vibrations. [12:28.120 --> 12:29.120] Here's what we learned. [12:29.420 --> 12:40.220] We did not know enough about vocal sounds to progress further, given the limited amount of time we spent on the Pratt software, and the pitch analysis idea was put on hold while we explored rhythm analysis. [12:42.920 --> 12:46.260] So here's the output of the first pitch analyzer. [12:46.580 --> 13:01.340] This pitch analyzer script maps a pitch range of a four-string violin, once again, from G3 to E7, and returns statistics about the input sample in a table of statistics, and the visualization outputs show two charts. [13:01.880 --> 13:06.820] One displays the waveform, and the other displays the pitch contour from the audio sample. [13:07.400 --> 13:12.380] We could use this to analyze audio samples and transcribe recordings into musical notation. [13:14.920 --> 13:17.100] And here's the second pitch analyzer. [13:17.700 --> 13:22.660] The second pitch analyzer script returns a pandas data frame storing the analysis results. [13:23.400 --> 13:38.920] After reviewing the script with Andrew, we weren't sure what those F1, F2, F3 jitter and shimmer values were, what they meant, what they represented, or how they could actually help a student work on intonation. [13:40.060 --> 13:42.620] So we went digging into the Pratt documentation. [13:43.260 --> 13:50.100] First, we found that the F1, F2, F3 results refer to something known as formants. [13:50.960 --> 13:54.580] Formants F1 and F2 are related to vowel height and vowel place. [13:55.680 --> 14:04.280] This expands our previous focus on pitch and frequency to include vocal height, sorry, to include vowel height, vowel place, and formants. [14:05.040 --> 14:10.660] But the question remains, how does this apply to violin practice? [14:12.900 --> 14:14.680] We weren't so sure about jitter either. [14:15.260 --> 14:21.900] And it turns out that jitter values beyond a certain threshold is associated with pathology, according to something known as the MDVP. [14:22.940 --> 14:25.140] And at this point, we were deep in the weeds. [14:26.460 --> 14:36.140] Shimmer values are also defined in Pratt software, and they represent amplitude variations in vocal fold vibrations, which is a key indicator of acoustic voice quality. [14:38.140 --> 14:40.960] This isn't what we expected when we first started our project. [14:41.240 --> 14:44.080] Yet, it's an important clue linking voice analysis to the violin. [14:44.760 --> 14:47.940] We did not know how this would be meaningful to violin students, though. [14:49.020 --> 14:51.220] So it was time to find another path. [14:51.700 --> 14:53.600] Next, we explore rhythm analysis. [14:55.440 --> 15:01.820] After moving away from pitch analysis, we pivoted to rhythm analysis, and we created a rhythm analyzer. [15:02.800 --> 15:08.780] This is a command line utility for rhythm analysis that estimates the tempo of a recording in beats per minute. [15:09.160 --> 15:16.700] And the goal here was to review a student's recording note length for a particular exercise when working on improving timing. [15:17.740 --> 15:21.280] The basic idea here is that we want to know if a student's playing in time. [15:22.680 --> 15:29.540] We prompted LM studio to draft the rhythm analysis script using the library. [15:31.540 --> 15:44.220] And for the code review part, after Andrew and I had a chat to develop the idea of the rhythm analyzer further, we recognized there would be more factors needed for an effective rhythm analyzer than just beats per minute. [15:45.100 --> 15:53.640] For instance, we noticed we didn't take rubato into account, and we would need a way to align a tempi of multiple recordings for the analysis tool. [15:55.040 --> 16:05.680] We learned that while rhythm analysis was a harder problem to solve than at first glance, it led to our exploration of the Librosa library documentation and MFCCs. [16:06.420 --> 16:13.500] At this point in mid-July, we also discovered there were many apps and software features available for pitch and rhythm analysis. [16:14.220 --> 16:19.880] So we thought to try something new that hasn't been addressed by existing tools after encountering MFCCs. [16:21.020 --> 16:27.020] We didn't know what MFCCs were and researched the applications, then found a breakthrough. [16:30.040 --> 16:33.140] And here is the breakthrough itself. [16:33.360 --> 16:42.400] We have a quotation here that says, antique Italian violins generally resemble basses and baritones, but Stradivari violins are closer to tenors and altos. [16:43.920 --> 16:52.460] If you recall from HOPE 15, Hack the Violin, Andrew mentioned singing a phrase and then playing it on the violin. [16:53.560 --> 16:56.300] So what's the connection between singing and the violin? [16:57.760 --> 17:11.940] Turns out, while learning about MFCCs and their applications in voice and speech recognition, we found a key piece of information that both validates our pitch analyzers and strengthens the idea of violins sounding like the human voice. [17:12.660 --> 17:29.020] In their 2018 publication, Acoustic Evolution of Old Italian Violins from Amati to Stradivari, Huang Qingtai and other researchers used Pratt software to analyze antique violin audio, as well as audio from male and female singers. [17:29.680 --> 17:36.940] They found that voice-like quality of violins aligns well with the rise of professional female singers. [17:37.460 --> 17:48.360] And the idea of the violin sounding like the human voice goes further back to 1751 when Francisco Gemignani published The Art of Playing the Violin. [17:49.440 --> 17:54.800] And so, perhaps we could try analyzing violin sounds with MFCCs. [17:56.620 --> 17:58.540] This led to our tamper analyzer. [17:59.220 --> 18:03.920] After reviewing use cases of MFCCs, we believe they could apply to our violin project. [18:04.680 --> 18:09.540] And this tamper analyzer is a command line utility built with the Librosa Python library. [18:09.960 --> 18:16.920] It processes a WAV file or multiple WAV files as the input and generates a tamper profile of the audio sample. [18:17.700 --> 18:32.240] So, during our code review, we noticed that the output of this analysis script includes results for 13 MFCC values describing their perceived tamper qualities, a CSV file, and a dashboard. [18:33.980 --> 18:39.820] We learned that we didn't quite understand how to interpret the resulting MFCC values or dashboards just yet. [18:40.180 --> 18:44.460] So, we focused on the detailed MFCC coefficient analysis results. [18:44.880 --> 18:51.240] And at this point, we weren't sure how that analyzer's results could be meaningful to a student or a teacher. [18:53.040 --> 18:55.400] So, here's some of the output. [18:55.740 --> 19:03.340] This is the result of the tamper analyzer after running it on a wave file sample of a G major two-octave scale recording. [19:06.990 --> 19:17.210] Notice how there's an overall tamper profile with descriptors like brightness, harmonic richness, attack character, warmth, clarity, and so on. [19:18.510 --> 19:22.350] These words help us describe how we hear and perceive tamper. [19:25.230 --> 19:26.430] Can I make this full screen? [19:35.970 --> 19:42.830] So, I won't be able to navigate the screen directions through here, if I do that. [19:48.940 --> 19:49.680] Sorry about that. [19:52.500 --> 19:54.880] Is it also blurry from your perspective? [19:55.780 --> 19:57.420] Yeah, it's pretty hard to read those things. [19:57.420 --> 19:57.520] Yes. [19:58.140 --> 19:58.700] Okay. [20:01.380 --> 20:04.220] Let's see if I can adjust that in a second. [20:07.120 --> 20:08.360] Brightness, built-in display. [20:18.730 --> 20:19.590] Does this help? [20:22.050 --> 20:23.710] They do come up again, so... [20:27.190 --> 20:28.670] Is that better? [20:29.710 --> 20:30.050] No. [20:30.470 --> 20:30.710] No? [20:30.990 --> 20:31.230] Okay. [20:37.340 --> 20:40.460] Andrew, what are your thoughts on putting the slides online? [20:41.180 --> 20:42.320] Yeah, that's all good. [20:42.460 --> 20:44.240] And we can definitely do that. [20:44.420 --> 20:47.720] And at this point, this is just to give you an idea of what we were looking at. [20:48.000 --> 20:52.200] And I think we responded to it the way that you all probably are. [20:52.320 --> 20:53.680] What's this large amount of text? [20:54.180 --> 20:58.920] But the bit at the end has the keywords, which sort of perked up on our minds. [20:59.140 --> 21:04.920] So, maybe show them the graph, and then we can carry on as long as they can see what's going on generally. [21:10.440 --> 21:11.780] So, the next slide. [21:13.220 --> 21:14.120] Okay, let's see. [21:16.260 --> 21:17.060] Okay, cool. [21:17.160 --> 21:18.980] I got the slide control back. [21:21.780 --> 21:24.580] This didn't mean anything at all at the time. [21:25.920 --> 21:30.680] We did not contribute to the school shortage, as reported from the last talk, though. [21:34.620 --> 21:37.060] So, this is the Timbre Analyzer dashboard. [21:38.080 --> 21:42.900] At this point, we sat together in a meeting thinking, well, what do we do with this exactly? [21:44.220 --> 21:51.020] And giving the Timbre Analyzer WAV files means getting a lot of data back, which felt far from artistic expression. [21:51.280 --> 21:54.320] And it did not appear to make Learning Violent any easier. [22:00.740 --> 22:01.420] Okay. [22:01.640 --> 22:02.920] So, what did we do next? [22:05.300 --> 22:05.980] Okay. [22:05.980 --> 22:06.900] Shall I go? [22:08.160 --> 22:09.100] Yeah, sure. [22:09.820 --> 22:10.580] All right. [22:10.800 --> 22:16.720] So, what we did next, using AI for comparative analysis and a breakthrough. [22:17.020 --> 22:27.820] So, here we are thinking that we have Timbre Texture Analysis, which looks interesting in that it is using some musical expressive language, just at the end there. [22:28.020 --> 22:29.120] But what does it mean? [22:29.620 --> 22:46.040] So, at the time, in the meeting, I played an identical set of notes to EVMBAT's WAV file example, but I played it in a highly, sort of, exaggerated, expressive way, as opposed to the plain, very correct performance that we already had as a WAV file. [22:46.260 --> 22:51.820] Because we wanted to see if there would be any difference, really, between the two MFCCs analysis. [22:52.320 --> 22:58.020] I mean, there was... is it just the sound of a violin is the same? [22:58.320 --> 23:01.740] And there isn't much variety, no matter what you do with the violins. [23:02.900 --> 23:08.900] So, there was a little wording difference and a little change in the numbers, but we still did not know what it meant. [23:09.040 --> 23:16.400] So, we asked AI, Chachupiti, to compare the two MFCC results. [23:16.960 --> 23:18.480] And this is what we got. [23:18.660 --> 23:19.160] Next slide. [23:21.700 --> 23:22.260] Okay. [23:23.040 --> 23:28.120] Now, what jumped out immediately... and there's more, so it'll be clearer. [23:28.560 --> 23:38.680] There's smaller versions of this to see... is vocabulary, softer, quieter, left treble heavy, smoother, richer, harmonic content, resonant, warm. [23:38.780 --> 23:50.180] If you jump right down to 11 and 12, we get much less nuance, more plain or generic sounding, less distinctive, more generic or synthetic feeling timbre. [23:50.820 --> 23:58.540] So, that file that was plain was played that way deliberately, and the exaggerated one was played that way deliberately. [23:58.720 --> 24:06.260] But Chachup can hear the difference, or the MFC analysis shows the difference, and Chachup was able to interpret it. [24:06.400 --> 24:08.060] So, just to give a closer look. [24:09.420 --> 24:10.160] Next slide. [24:12.440 --> 24:14.140] It gets into a little detail. [24:14.420 --> 24:17.080] Both are soft, but the second file is even quieter. [24:17.620 --> 24:19.280] Both files are very bright, but slightly. [24:19.580 --> 24:21.520] But the first is slightly brighter and more brilliant. [24:21.660 --> 24:24.680] And I'll skip further down. [24:25.040 --> 24:26.080] Timbral nuance. [24:26.440 --> 24:28.740] First file is more expressive and detailed. [24:29.280 --> 24:31.080] Second file is more uniform polish. [24:31.340 --> 24:34.660] Now, this is exactly our opinion of these two files. [24:35.720 --> 24:38.300] So, first file, more sparkle top end. [24:39.280 --> 24:41.280] Second, richer overtone series. [24:41.860 --> 24:42.300] A little warmth. [24:42.500 --> 24:43.920] Slightly more body in the mid-range. [24:44.640 --> 24:45.920] And so on and so on. [24:45.960 --> 24:47.200] You get to the overall feel. [24:47.420 --> 24:50.040] First file, artistic nuance, slightly brighter. [24:50.280 --> 24:52.480] Second file, clean and resonant, but more generic. [24:53.020 --> 24:59.200] This is literally how we would describe these two instances of audio. [24:59.560 --> 25:17.480] So, here we have an MFCC analysis, and then a translation using the language that they use to describe the MFCCs, with Chachup making an effective and highly accurate, really, response. [25:17.700 --> 25:18.140] Next slide. [25:21.780 --> 25:22.280] Sorry. [25:22.520 --> 25:22.980] Next slide. [25:24.380 --> 25:26.620] So, that's very, very exciting. [25:26.960 --> 25:40.180] Like, it's a very exciting moment, whereby we've taken data and we've turned it into translatable words that you could then, you know, use to describe sound and also say to someone else. [25:40.980 --> 25:43.940] Then we thought, okay, let's take this a little further. [25:44.260 --> 25:51.980] I uploaded all my notes from my talk last year and all my teaching notes, because I had previously compiled them into one file. [25:52.380 --> 26:01.340] And in the early days, when with fights and Langchain, tried talking to my file, my notes, to see what kind of response I would get. [26:01.700 --> 26:04.140] I had about an 80% nice response at the time. [26:04.380 --> 26:06.000] But chat's much more advanced. [26:06.140 --> 26:07.080] Now it's two years later. [26:07.860 --> 26:09.000] Gave them that file. [26:09.240 --> 26:20.040] And then we asked chat to figure out what would you tell the playing player in order to get them to sound like the more artistic nuance player. [26:20.420 --> 26:22.960] So, next slide. [26:26.260 --> 26:35.300] So, it actually went into the notes and found exactly what I would have said through my teaching. [26:35.540 --> 26:37.960] So, the number one, try to sing what you're playing. [26:38.120 --> 26:39.400] You must know what you're trying to express. [26:39.660 --> 26:45.660] That is literally the number one biggest hack to playing the violin and learning it quickly and easily. [26:46.160 --> 26:46.420] Right? [26:46.920 --> 26:48.240] And then it makes some insight. [26:48.540 --> 26:53.360] The second file lacks expression, suggests the musical concept may not have been fully internalized. [26:53.540 --> 26:58.240] And the suggestion of action before playing, sing the entire melody phrase by phrase. [26:58.440 --> 27:01.860] Let the feel of the music drive your phrasing, not just the notes. [27:02.100 --> 27:04.100] So, chat picked out the right part of the text. [27:04.280 --> 27:04.800] Next slide. [27:06.880 --> 27:07.980] Do the sound. [27:08.180 --> 27:09.160] You're doing sound. [27:09.380 --> 27:12.440] The hands act as a result of the effort to do the music. [27:13.620 --> 27:16.760] When you play, it ain't the sound you want to hear. [27:16.860 --> 27:19.740] Again, this is right out of the text, right out of the hack. [27:19.940 --> 27:20.480] Next slide. [27:21.280 --> 27:24.400] And it even got into some of the more practical points. [27:24.580 --> 27:27.080] So, using the left hand grip to shape tone. [27:27.300 --> 27:29.020] Grip retraction, ultimately the fingertip. [27:29.300 --> 27:31.040] The grip actually keeps you there. [27:31.200 --> 27:34.740] And you do transfer the expression through touch. [27:34.920 --> 27:38.780] And you do so in your left hand, also in your right hand. [27:39.080 --> 27:47.420] But the nice one I liked was at the bottom of the slide there, you say, it says, imagine petting a dog to express the motion through touch. [27:47.640 --> 27:51.060] That's an analogy always used, especially with young people. [27:51.720 --> 27:53.100] Nothing complex about it. [27:53.180 --> 27:55.660] Everyone can pet a dog and tune into their sets of touch. [27:55.880 --> 27:56.620] Next slide. [27:57.500 --> 27:58.540] Aiming with your ears. [27:58.720 --> 27:59.600] Hear what you want to hear. [27:59.940 --> 28:01.080] before you play it. [28:01.800 --> 28:03.380] Prehear the notes tone. [28:03.560 --> 28:04.080] Next slide. [28:05.620 --> 28:07.240] And it added the bow part. [28:07.360 --> 28:08.720] The bow is where expression happens. [28:09.180 --> 28:11.660] Balance the bow on the string and feel contact. [28:12.840 --> 28:13.740] Next slide. [28:15.560 --> 28:17.980] And it tagged mental preparation. [28:18.160 --> 28:19.380] Be ready before you play. [28:19.640 --> 28:22.360] Have an idea of the emotion and sound you want. [28:24.580 --> 28:26.400] Then we get to the final summary. [28:27.220 --> 28:28.100] Next slide. [28:29.500 --> 28:32.060] So you've got a nice summary here from chat. [28:32.660 --> 28:35.880] Based on those notes, telling the player what to do. [28:36.020 --> 28:38.720] And it's fantastically useful advice. [28:38.960 --> 28:41.420] It's entirely on point, based on the notes. [28:42.360 --> 28:44.160] And chat did all the heavy lifting. [28:44.640 --> 28:53.400] Now, if you've been working, teaching for like, you know, 30 plus years or even more, you'll pull these things out and you'll say them. [28:53.580 --> 28:57.440] But if you're starting out, or if you're by yourself at home, you won't. [28:57.560 --> 28:59.360] So it's a fantastic tool. [28:59.820 --> 29:00.320] Next slide. [29:01.040 --> 29:03.580] So why stop when you're having fun. [29:03.820 --> 29:07.080] We got this amazing amount of information from that. [29:07.240 --> 29:10.560] We thought, well, let's give it some of the historical treatises. [29:10.880 --> 29:13.300] So the Geminiah treatise, which E.P.M. [29:13.320 --> 29:14.360] referenced before. [29:14.620 --> 29:15.740] Had a copy of that. [29:15.980 --> 29:17.040] So uploaded that. [29:17.560 --> 29:19.560] And it gave the full spectrum. [29:19.720 --> 29:21.880] But I'm just... I've got the summaries in the slides here. [29:22.080 --> 29:26.380] So you can see expression of the soul, judgment of the air, 18th century language. [29:26.960 --> 29:29.940] Pre-hear the tone before playing, adjust fingers bone to match. [29:29.940 --> 29:32.500] Taste over technique, emotion. [29:32.940 --> 29:44.140] These are all similar concepts, as Geminiah would have expressed it, to explain to the person who played the generic file, the generic sound rather, how to make a more expressive sound. [29:45.060 --> 29:46.620] And then we just piled on. [29:46.740 --> 29:47.280] Next slide. [29:48.360 --> 29:52.920] Uploaded a French version of the flesh, called Fleshes, treatise on the violin. [29:53.700 --> 29:55.560] And language is no issue for chat if you teach. [29:55.980 --> 29:58.720] And you can see similar advice here. [29:58.960 --> 30:00.380] To tone must be colored. [30:00.660 --> 30:01.620] Legato must breathe. [30:02.060 --> 30:03.620] The mind precedes the hand. [30:04.040 --> 30:08.440] To make your violin sing, you must hear, you must first hear its song in your head. [30:08.920 --> 30:09.580] Next slide. [30:10.760 --> 30:13.620] And then Galamian, very famous violin teacher. [30:13.980 --> 30:18.100] You can see he also has interesting advice. [30:18.300 --> 30:19.280] They're all a little different. [30:19.780 --> 30:27.020] But basically, and this is a good level in a way too, because we could ask chat for a deeper dive into any one of these. [30:27.320 --> 30:29.700] They could pull more phrases out of country pieces. [30:30.340 --> 30:32.920] Or give us the exact spot to go and look at the course. [30:33.680 --> 30:34.420] Next slide. [30:35.560 --> 30:38.640] We checked out what Francis Gatti would have to say. [30:39.280 --> 30:40.620] Particularly ninth phrase. [30:41.000 --> 30:45.600] If the player fought, you can think like a singer, phrase like a poet, and bow like a dancer. [30:46.220 --> 30:46.480] Okay. [30:47.200 --> 30:48.100] Next slide. [30:48.940 --> 30:51.100] Leopold Auer, famous Russian teacher. [30:52.920 --> 30:54.400] And then Mozart's dad. [30:54.600 --> 30:55.160] Next slide. [30:56.760 --> 30:59.280] Who was really into the violin. [30:59.840 --> 31:04.020] And apparently constantly nagged his son to do more violin playing. [31:04.840 --> 31:06.940] But this is a very old practice. [31:07.180 --> 31:09.120] But let your bows speak like a tongue. [31:09.320 --> 31:10.100] Always like that one. [31:10.340 --> 31:13.580] Your fingers act like a poet, and your ears judge like a prince. [31:13.860 --> 31:19.080] So, I mean, here we are with these pieces. [31:19.340 --> 31:19.840] Next slide. [31:20.400 --> 31:24.280] And I think the thing to keep in mind is we got this all right away. [31:24.980 --> 31:30.340] You know, I'm sure everyone here has experienced chatubiti, so they understand what I mean by that. [31:30.580 --> 31:34.680] It's like we popped it in, we asked our stuff, we got our results immediately. [31:36.400 --> 31:39.280] Which is the heavy lifting being done, for sure. [31:39.660 --> 31:43.520] I do remember reading all those statistics in university, and it took a while. [31:44.320 --> 31:48.260] So, just to add, I did play some more complex longer examples. [31:49.440 --> 31:51.840] And you get the same type of results. [31:52.440 --> 31:54.540] And so, AI is doing heavy lifting. [31:54.820 --> 32:08.020] We can record a piece, thanks to that amazing script, do a quick tabular analysis, submit the analysis, and a comparison analysis, draw on any treatise, or set of parameters we provide to get the immediate practical. [32:08.860 --> 32:12.520] So, at this point, we also wonder, well, why did it work so well? [32:13.420 --> 32:16.920] And what are MFCCs, anyway? [32:17.220 --> 32:19.560] And I'll pass you back to Evie and that. [32:20.020 --> 32:20.740] Evie Smith- Thanks, Andrew. [32:24.070 --> 32:26.570] So, the slide's about Labrosa and MFCCs. [32:27.550 --> 32:38.210] After the surprisingly emotive and expressive interpretations of the timbre analysis output, we revisited the ROSA and MFCCs to better understand both of these topics. [32:38.570 --> 32:47.670] And what we'll do next is we'll dive into MFCCs, the MEL scale, and why MFCCs are powerful for understanding expression in sound. [32:49.250 --> 33:00.250] MFCCs, also known as MEL frequency septual coefficients, are widely used for voice recognition, music genre classification, and music instrument identification purposes. [33:01.510 --> 33:09.490] MFCCs are a set of numerical values that describe spectral characteristics of sound for machine learning applications. [33:09.930 --> 33:20.090] They were developed in 1980 by Paul Mermelstein and Stephen Davis in their research on parametric representations of acoustic data in speech recognition systems. [33:20.790 --> 33:23.650] MFCCs are measured in MEL scale units. [33:24.050 --> 33:33.990] The MEL scale, used in MMCC computation, splits sound into different frequency bands with more attention given to frequencies used to understand human speech. [33:35.030 --> 33:52.830] The MEL scale frequency bands are spaced logarithmically above a thousand hertz, compressing the sound, and linearly below a thousand hertz to align with how humans perceive sound based on a psychoacoustic analysis research from the 1930s and 1940s. [33:53.150 --> 33:54.650] So what does that mean? [33:55.970 --> 34:01.990] MFCCs represent correlations to how humans perceive sound as a subjective experience. [34:02.970 --> 34:18.150] MFCCs tie together the human voice analysis, singing and violin concepts, and MFCCs align with how we perceive sound, capturing the shape of sound to infer characteristics of timbre and texture and fit our use case quite well. [34:20.410 --> 34:28.190] So this table here shows the 13 MFCCs used in the timbre analyzer to describe timbre and textural qualities of sound. [34:29.050 --> 34:51.530] We'll have some examples here, such as the brightness, sharpness, harmonics, attack, body and warmth, clarity, whether a sound is more woody or nasal, brilliance, airiness, texture, timbre detail, and character, which is super helpful for identifying small nuances in the sound. [34:53.910 --> 34:55.430] So what does this mean? [34:58.150 --> 34:59.910] This is a little bit of a reflection. [35:00.930 --> 35:05.550] We didn't know that this project would result in a useful tool for violinists. [35:06.230 --> 35:17.290] And further, the details starting from coding the scripts in our research phase helped us build up from a technical and analytical perspective to an expressive and human-centered teaching approach. [35:18.270 --> 35:30.030] When we first relied heavily on numerical results and statistics, one of my concerns was that this project would remove human creativity, emotion, and feeling from violin practice. [35:30.550 --> 35:35.490] Luckily, with Andrew's experience and insight, the technical details did not overshadow the project. [35:36.730 --> 35:40.630] And as a violin student, I felt that this was an awesome resource. [35:41.130 --> 35:47.910] Being able to review feedback from Andrew's teaching notes helps practice and it helps develop habits and routines. [35:48.730 --> 36:01.410] Applying references from violin treatises created a panel-style feedback from many instructors at once, which is incredible for learning and growing as a violinist, both technically and musically. [36:01.890 --> 36:03.890] Now, I'm handing this to Andrew. [36:05.490 --> 36:06.450] Thank you, Ibi and Beth. [36:07.150 --> 36:16.110] So, we found a way to marry violin playing in AI, which seems very simple and even obvious now. [36:16.330 --> 36:24.390] I just want to point out that we did not find any other versions or suggestions for the pipeline we created prior to discovering it. [36:25.010 --> 36:32.670] And particularly, Ibi and Beth's script for putting together the MFCC analysis for us. [36:33.110 --> 36:38.830] Now, I did ask chat last night if there was any other study that had done this, because I was curious at this point. [36:39.710 --> 36:47.950] And they said, MFCCs plus violin current state, MFCCs are already used in research to measure and compare violin timbre. [36:48.170 --> 36:54.010] For example, distinguishing between instruments, visualizing tone quality changes, or classifying performances. [36:54.350 --> 37:01.250] But they said, what's missing is there's no widely documented system that takes MFCC results. [37:01.590 --> 37:06.150] Did we actually say there were mal-frequency Seppström coefficients? [37:06.470 --> 37:09.730] Anyway, that's what MFCC, I think you must have said it, right? [37:10.090 --> 37:13.250] Results, and translates them into actionable violin technique advice. [37:13.630 --> 37:16.810] And we certainly did not find anything when we were looking. [37:17.070 --> 37:21.870] So, as far as we know, that's a first step in this. [37:22.470 --> 37:24.530] Currently, that translation is manual. [37:24.690 --> 37:30.890] The teacher listens, interprets, and applies historical technical knowledge without any real help other than experience. [37:31.290 --> 37:35.170] So, but what we have shown you today bridges that gap using AI. [37:35.870 --> 37:51.170] And as chat would say, this would be the first pipeline that joins objective signal analysis with subjective historically informed coaching, effectively turning spectral data into personalized tradition-based violin instruction. [37:51.770 --> 37:53.390] It's a mouthful, but anyway. [37:53.830 --> 38:00.090] So, our next step might be create an agent, as I'm sure you would imagine right away. [38:00.530 --> 38:14.630] Create an AI virtual teacher built from maybe quotations and instructions from any violin master or any violinist who's active, or a combination of them, combined with MFCCs of their performances. [38:15.150 --> 38:26.010] So, we would end up with the ability to give specific player-tailored feedback in the style of that master or combination of masters to help a player change their playing to that style. [38:26.390 --> 38:39.130] One could essentially have a lesson with any great player with far greater depth than previously available from reading the treatise or method book, or listening to an interview, or even watching a master class. [38:39.310 --> 38:45.670] It's the nearest you could get to having an experience or nearer than being, you know, obviously. [38:47.350 --> 38:58.790] So, ultimately, this could expand even further in terms of discovering a best learning style of visuals and tailoring the musical advice and vocabulary for those particular styles. [38:58.930 --> 39:01.170] I mean, that's what all teachers do. [39:01.350 --> 39:05.790] They figure out what to say to this particular person at this particular stage. [39:06.130 --> 39:14.830] And I am sure there are further applications that will arise, and I really look forward to people building on this basically initial experiment. [39:15.550 --> 39:40.190] Anyway, so at this point, I want to thank you there to EVMBAT for your amazing work, finding MFCCs and writing up the code for the project, and, you know, your feedback on all the violin work, the smooth code, and just for being a fantastic human to work with and discover new things with. [39:40.550 --> 39:42.150] It was a great experience. [39:43.670 --> 39:49.370] And I want to thank the HOPE conference, because this is such a great thing. [39:49.510 --> 39:52.290] I'm really sorry I was unable to be there in person this year. [39:53.290 --> 39:56.290] And particularly, I want to thank Lindsey for throwing my ticket. [39:56.950 --> 40:02.670] And thanks to 2600, who have been a beacon of sanity for so long. [40:02.790 --> 40:07.730] And I do think of 2600 as one of the most important entities in America. [40:08.030 --> 40:10.750] And I'm grateful for their continued efforts. [40:11.850 --> 40:21.050] And maybe most importantly, special thanks to all of you for listening to our talk and making this experience even more meaningful. [40:21.230 --> 40:22.510] Very much appreciated. [40:22.770 --> 40:25.330] And I'll just hand you back now to EVMBAT. [40:27.110 --> 40:28.010] Thanks, Andrew. [40:28.230 --> 40:31.650] And thank you for this opportunity to collaborate on a violin project. [40:31.850 --> 40:38.850] I could not have gotten as far along as I did without your direction and experience in both teaching and AI. [40:39.710 --> 40:46.050] Thanks to the HOPE 16 organizers, the audience here in person, as well as those of you tuning in online. [40:46.390 --> 40:49.630] We hope you enjoyed this talk today, and have a great time at the conference. [41:01.060 --> 41:02.440] Does anyone have any questions? [41:09.860 --> 41:10.320] Hi. [41:10.500 --> 41:11.160] I have a handful. [41:11.520 --> 41:12.740] The short ones first. [41:13.000 --> 41:13.880] Is any of this on GitHub? [41:14.780 --> 41:16.740] Is any of the code on GitHub? [41:17.680 --> 41:18.980] Not yet. [41:18.980 --> 41:20.960] We do have the source code, though. [41:21.240 --> 41:21.480] Cool. [41:21.680 --> 41:24.000] If it was, I would love to look over what you've got. [41:24.500 --> 41:26.820] Have you tried other instruments or just violin? [41:27.440 --> 41:28.680] Have we tried other instruments? [41:30.120 --> 41:35.220] There was a recording... Andrew, what was that guitar with the metal piece again? [41:35.920 --> 41:37.420] We recorded it, but we didn't. [41:37.820 --> 41:39.340] I played around with it. [41:39.480 --> 41:41.020] It's a good point. [41:41.180 --> 41:52.020] I was thinking, I imagine that wind instruments, well, would probably do really well with this, just because of their ranges and, again, a similarity to human voice. [41:52.340 --> 41:58.100] But it would definitely be worth... I'm inclined to try every single instrument and see what you get. [41:58.400 --> 42:06.680] Because I suspect that even if it wasn't a really good match to the human voice in terms of range and things like that, that you can nonetheless adapt the code. [42:06.840 --> 42:06.940] Yeah. [42:07.080 --> 42:08.900] That would be a great way to go forward. [42:09.080 --> 42:09.120] Yeah. [42:09.420 --> 42:12.900] Does it deal with things like double stops and, like, pizzicato chords very well? [42:13.020 --> 42:16.860] Or since the human voice doesn't do that so well, it falls apart there? [42:19.160 --> 42:22.760] I did play some Bach into it. [42:22.940 --> 42:27.300] And I got analysis back, which was pretty good. [42:27.600 --> 42:36.940] I mean, I think what I'd have to do there is take a section that was particularly polyphonic and then see if I could get more detail. [42:38.100 --> 42:48.540] But because it's expressive qualities, I don't think that's really going to make a lot of difference because it's analyzing sort of a combined frequency and how expressive is it. [42:48.660 --> 42:51.720] So it's not zeroing in on it being any one note. [42:51.840 --> 42:57.600] And if you think about the sound of the human voice, I mean, that's a rich, almost polyphonic sound. [42:57.740 --> 43:00.840] I mean, we think of it as only one pitch when we would analyze the pitch. [43:00.980 --> 43:05.340] But really, I mean, the resonant sound of a voice is just full of all kinds of frequencies. [43:05.640 --> 43:08.860] So I imagine I would have to verify that that would work okay. [43:09.220 --> 43:09.340] Sorry. [43:09.620 --> 43:11.840] And finally, sorry to ask so many here. [43:11.940 --> 43:13.280] I could talk to you for hours on this. [43:13.760 --> 43:14.040] That's good. [43:14.040 --> 43:22.400] Does it always require a comparative sample of two different, like here's the one that we're going for and here's the student? [43:22.660 --> 43:27.420] Or does it simply, does the MFCC process simply analyze one standalone? [43:28.920 --> 43:30.900] Sorry, I didn't quite catch the... [43:30.900 --> 43:35.780] Does it require a comparative sample of two, like a student and a teacher? [43:37.020 --> 43:41.340] Or does it analyze like a singular performance on its own? [43:42.960 --> 43:45.520] So we use the comparative sample. [43:45.700 --> 43:56.420] And the reason that we had, anyway, to use the comparative sample is I thought that if it tried to give an absolute answer, that's not really going to work out. [43:56.540 --> 44:07.220] I mean, essentially, and it's just my opinion, anybody who is listening to somebody and giving feedback is comparing it to something, even if it's just their idealized version of the way things are. [44:07.220 --> 44:14.480] And I did think about just trying to ask chat, see what they would come back with. [44:14.640 --> 44:17.680] But I don't think it's really meaningful without some kind of comparison. [44:17.760 --> 44:22.200] And I think you would always want to have some example for somebody to go off. [44:22.920 --> 44:26.000] But, you know, that's to be explored kind of thing. [44:26.520 --> 44:35.660] The nice thing about the example is, is say we wanted to have, you know, Hilary Hahn be our virtual teacher or make a virtual version of her. [44:35.840 --> 44:42.820] We can take all her recordings, we can analyze them all and get a very nice profile of what she's going to do. [44:42.960 --> 44:46.440] And it just seems like I can't see a way out of comparison at this point. [44:46.760 --> 44:48.560] So anyway, I know that helps. [44:51.100 --> 44:51.940] Good afternoon. [44:52.360 --> 44:55.260] I am a, uh, I'm an adult learner of the violin. [44:55.460 --> 44:57.100] I've been doing it about seven years now. [44:57.320 --> 45:13.280] And doing it outside the framework of, uh, you know, a school or college or some kind of conservatory or something like that means that I've had to struggle over the years with finding direction and identifying where to go with it next. [45:13.860 --> 45:19.140] Uh, are there any tools that you would find that are interesting to explore? [45:19.360 --> 45:28.940] Basically, I'm looking for sources of inspiration and different ways of approaching my learning that would motivate me and inspire me. [45:29.080 --> 45:31.760] Do you have any ideas or places to look for that? [45:33.520 --> 45:38.240] Uh, yeah, I mean, maybe we can both, uh, we can both answer that. [45:38.480 --> 45:48.700] Uh, I, I say, I would say, uh, have a listen to the talk from last year because there's quite a, there's quite a bit of useful information in that. [45:48.960 --> 45:53.760] Um, and certainly if you were to contact us, I could make that available to you as well. [45:54.020 --> 46:03.940] Um, but I think, um, uh, EVM has the, has some, um, places that you go to and, and where you study that look very interesting, so. [46:05.920 --> 46:13.100] Yeah, I've, uh, I've, I've been taking classes at the NYC violin studio and they accept adult students. [46:13.380 --> 46:23.040] There's, um, group classes and as well as, uh, focus on both classical and contemporary works for adult learner, adult learners. [46:23.260 --> 46:34.240] Um, when it comes to direction, I suppose it also depends on what you would like to participate and contribute to in terms of learning the instrument. [46:34.600 --> 46:48.980] Um, there's acoustic instruments, acoustic electric instruments, four strings, five strings, six strings, uh, lots of possibilities out there in terms of the kind of music you want to create with the violin. [46:50.660 --> 46:57.680] So, uh, in addition, um, I was just thinking, uh, some of the, uh, it depends on your style. [46:57.920 --> 47:14.360] I mean, the style that you're interested, but, but if it was classical, for example, if you follow Hilary Hahn on her social media, and if you, if you went through it, she did these 100 days practice, uh, routine where she was doing 100 days in a row and everyone would join and she'd post a lot of, [47:14.360 --> 47:17.160] of, uh, tips and she had goals. [47:17.440 --> 47:25.820] She, which is really very, very, well, she's a great artist, but also very, uh, helpful in terms of what she can translate, how she can translate that to other people. [47:25.920 --> 47:30.000] But if you were more into bluegrass, I would go find a mark of honor. [47:30.740 --> 47:42.380] I, I would go and find, there's, there's so much, uh, uh, information on YouTube and generally in social media that if you're looking at that, I would go look at those players and I'd listen to what they have to say. [47:42.620 --> 47:49.200] And a lot of times it's not like a detailed lesson plan that they're going to give you, but just a casual remark that might spark something for you. [47:49.340 --> 47:53.640] You know, you know, I always do this with my left hand and you're like, oh, I'll just try that. [47:53.640 --> 48:01.920] And a lot of times it's like one little thing opens a huge, uh, like opens a door to a huge advancement in your playing. [48:02.100 --> 48:05.640] So if you're looking at your inspiration, I would say it's a good time for that. [48:05.760 --> 48:07.580] You can go and find, go and find that. [48:07.720 --> 48:12.360] Follow the music because the music will take you to the type of player you want to listen to. [48:13.660 --> 48:17.380] Oh, and to, to add to that, uh, follow your musical taste. [48:19.760 --> 48:20.120] Yeah. [48:20.440 --> 48:21.060] Yeah, exactly. [48:21.320 --> 48:21.580] Exactly. [48:23.220 --> 48:25.680] And if you have a lot of musical taste, check it all out. [48:36.110 --> 48:36.650] Thank you. [48:36.970 --> 48:37.470] Thank you. [48:37.530 --> 48:38.150] Thank you very much.