you dont need to reproduce it, if you can finetune it.
The alignment in these models is really narrow, it doesnt take many finetuning steps to get out of the alignment basin.
the alignment is not data centric but done after the fact using RL...
You can also modify binaries. This doesn't address the benefits of reproduction more than a smidgeon. How can you answer "why the fuck does it work like this" if you can't inspect the process that built it? There's no replacement for "what was this built from?"
Of course, corporate/national investments need some moat, so I don't expect open source models to be competitive for a few years
Disclosure: it is a favor for my little sister.
I made it for her in 5min without much testing on recognition accuracy. Which was enough to blow her mind ;)
Given your feedback, i might do another loop on this.
Try ~crappy~ regular headset microphone and/or whistling for some basic edge cases :D
It worked quite well to start with, but then after a few bars deteriorated to mostly outputting pauses, and a stream of short notes (without pauses) when singing legato. Could be due to the aforementioned microphone, or is there some internal state that does not get cleared up when one clears up the score? I tried twiddling some of the variables that seemed related, but no noticeable improvement.
Please do. My kid has been messing with it but seems to be getting some notes wrong. Im sure some tweaking will make it behave correctly and I'd appreciate it.
reply