lowjieseng1810/malaysia-linguistics-lab
Credits + grant $3Live in productionDiscover the places, languages, voices and cultural stories that keep Malaysia's living languages alive.
Interactive learning and exploration platform for Malaysia's minority and indigenous languages.
- Python36.2%
- JavaScript26.6%
- CSS21.9%
- HTML15.2%
Placement
Every place and its price$3
$2.20 from converted credits; $0.80 granted by RepoRanker. Credits and grants are not card payments. This placement does not expire. Its rank holds until another repo spends more, and then this one moves down, never off. Taking the top of the board from here costs $15.
1 Review
Malaysia Linguistics Lab combines language learning, cultural exploration, and progress tracking into a substantial working application. It supports Iban, Kadazan-Dusun, Bidayuh, and Mah Meri through lessons, a searchable dictionary, quizzes, saved words, achievements, language comparison, community context, and a Three.js world explorer. The optional AI tutor stays separate from the core learning experience, so most features work without an external API. The Flask application includes local and Google authentication, CSRF protection, password hashing, login rate limiting, secure cookie options, production proxy handling, PostgreSQL support, and security headers. Password-reset tokens are hashed and expire after 30 minutes. The repository also includes focused regression tests for authentication, CSRF, failed-login limits, and achievement delivery.
The vocabulary work deserves specific credit. The repository contains a provenance ledger, identifies dataset licenses, preserves named dialect varieties, distinguishes historical Besisi material from modern Mah Meri, and explicitly rejects invented translations and sources without clear reuse rights. That discipline is important for a project involving indigenous and minority languages. The README provides a clear product walkthrough, screenshots, deployment guidance, known limitations, and an accurate explanation of which features require OpenAI.
The main opportunity is turning those good practices into an enforceable quality system. There is no CI workflow, so the existing tests are not run automatically. They also cover only a small part of a codebase containing roughly 20,000 lines of Python. The 7,000-line app.py holds routes, startup logic, configuration, and large course-data structures, which makes content review and behavioral testing harder than necessary. Moving course material into validated, versioned data files would let language experts review changes without navigating application code. The external CC BY and CC BY-SA vocabulary packs also need a dedicated data-license and attribution file because the repository-level MIT license does not explain how those datasets may be reused. Most importantly, the AI tutor is intentionally GPT-first and does not require retrieval from verified course material before answering. For low-resource languages, that creates a real hallucination risk that conflicts with the repository’s careful vocabulary standards. Translation and cultural answers should cite verified records, clearly label model-generated content, or decline when evidence is unavailable. Broader route, database, quiz, tutor, and vocabulary-validation tests would strengthen the platform considerably.
