armenian_llm_ecosystem
COPATeam/armenian_llm_ecosystem
Summary
This repository contains the complete reproduction code for the research paper "From Zero to Hero: An Open LLM Ecosystem for Armenian". It provides pipelines for building the ArmWeb corpus (4.37M Armenian news documents) and the ArmSTEM dataset (373K verified parallel STEM problems), along with training recipes for continued pretraining of LLMs (like Gemma) on Armenian data and evaluation harnesses.