展示HN:Vaghenu,一款了解音步的梵文颂歌朗读文本转语音工具
一个15年的梦想今天实现了。我开始攻读博士学位,梦想是创建一个能够完美吟唱任何梵文诗句的系统。
在这里,我发布了Vaghenu,这是一个具有韵律感知的梵文吟唱文本到语音(TTS)系统。这是世界上第一个关注韵律的开源梵文吟唱TTS。我将模型权重、训练脚本,甚至我精心收集的数据公开 - [https://prathosh.in/vagdhenu/](https://prathosh.in/vagdhenu/)
没有大型人工智能实验室,没有庞大的工程团队,也没有风险投资级别的预算。只有一位教授坚信人类最古老的知识传统之一值得拥有现代化的开放基础设施。
这个名字来源于《奥义书》中的一句话:“Vācaṃ dhenum upāsīta” - 像神话中的愿望成真之牛,Vāgdhenu旨在使梵文文本更易于学生、教师、研究人员和信徒们获取。
在这里测试实时演示,并告诉我你的意见 - [https://prathosh.in/vagdhenu/](https://prathosh.in/vagdhenu/)
整个系统,从数据收集到模型构建和演示,都是由一个人(我本人)使用我们在LatentForce构建的强大工具完成的。
我附上了系统生成的音频样本。
附言:这是我朋友的帖子,他们不在HN上。
查看原文
A 15-year-old dream has come true today. I started a PhD with the dream of creating a system that chants any Sanskrit shloka perfectly.<p>And here I am opening sourcing Vaghenu, a meter aware sloka-to-chant, TTS for Sanskrit . This is the world's first vrutta-aware, open-source TTS for Sanskrit Chanting. I am making the model weights, training scripts, and even data (that I meticulously collected) public - <a href="https://prathosh.in/vagdhenu/" rel="nofollow">https://prathosh.in/vagdhenu/</a><p>No large AI lab. No big engineering team. No venture-scale budget. Just a professor's conviction that one of humanity's oldest knowledge traditions deserves modern, open infrastructure.<p>The name comes from the Upanishadic phrase: "Vācaṃ dhenum upāsīta" - Like the mythical wish-fulfilling cow, Vāgdhenu is intended to make Sanskrit texts more accessible to students, teachers, researchers, and devotees everywhere.<p>Test out the live demo here and let me know your comments - <a href="https://prathosh.in/vagdhenu/" rel="nofollow">https://prathosh.in/vagdhenu/</a><p>The entire system, from data collection to model building and demos, is built by a single person (your truly) using the powerful harness that we are building at LatentForce.<p>I have attached a sample audio file generated by the system.<p>P.S: Posting on behalf of my friend, their aren't on HN.