Hello World!! My name is Donal Muolhoi and i trained a HmarBERT (maybe)!!
A brief pre-intro : My language (Hmar) has no digital footprint and no presence. Anything out there about our people is mostly wrong. So i decided even if i can’t code the least i could do is collect stuff so people can use it to train AI and i just started doing it. I had no concept of repositories then and i thought pdf files were the height of file formats, so i started collecting pdf files and up them to Notion. I thought this was enough.
Although slow i did manage to collect some books, i learned how to scrape websites, and i learned about json and csv and other data formats and Github and Hugging Face. But somewhere along the way i started feeling judged even though nobody had said anything, nobody even knew what i was doing. I felt like i was pretending to be smarter than i actually am but i can’t exactly go around advertising that i don’t know what i’m doing everytime i have to share my insights and opinions and stuff.
Come 2026, i started up some of the stuff i had scraped and i though i might just give up cuz nobody sees it anyways but i didn’t . I renamed my project to Hmar Heritage Foundation. I started reaching out to people and it is doing much better than it was a year ago. Of all the people i reached out to, a professor at a certain university suggested its worth attempting to train a Hmar model from a Mizo one and i thought about it and i did. I asked ChatGpt and i learned the technology is so matured that all i really needed was to have the courage. Of course its so much more than running a notebook on Colab or Kaggle but for the rest of us who the technology is intended for, its absolutely enough.
You can check out the models here : HmarBERT Arena - a Hugging Face Space by azinamotoe
I’m not sure if they mean anything but it’s still nice to think i created something that can generate text in my language.
I have learned alot of things about languages and my language in particular. Even though i still can’t code and probably never will (thanks to AI) and i haven’t really achieved my goals. I am thankful to developers, open source communities, and everybody else in between.
This is an invitation for anybody to check out our datasets at : hmar-heritage-org (Hmar Heritage Foundation) offer advice on what we should do next and how we can structure datasets. Its more than just one of me now. I have a friend who is just as uniformed as me and almost 10 on our WhatsApp group!! The only issue is when AI writes these dataset README’s it makes them sound overly academic which i then have to prompt it to tone it down. I don’t know how all this will be recieved but i still have to leave it out there anyways!!