JEMARI — Real-Time BISINDO Sign Translator
Translates Indonesian Sign Language (BISINDO) fingerspelling from a webcam into text and speech in real time with a CNN + Transformer.
- Role
- ML Engineer (research implementation)
- Client
- Felix Aria & Muhammad Raysar Al-Fatih — SMAN 1 Balikpapan (OPSI 2026)
- Year
- 2026
- Technologies
- PythonPyTorchMediaPipeFastAPIJavaScript
Overview
A working implementation of an OPSI 2026 research proposal: the full pipeline from dataset preprocessing and GPU training to a real-time web app.
The Problem
Digital communication remains hard to access for BISINDO signers; an automatic translator running directly from a camera was needed.
Approach
Hands are detected and cropped with MediaPipe; MobileNetV3-Small extracts per-frame features, fused with 21 hand-joint landmarks, and a Transformer encoder models 8 consecutive frames before classifying into 26 letters.
Challenges
The first version reported 96.99% accuracy, which turned out to be invalid: 56.9% of the dataset was offline augmentation, so the test set contained copies of training images. Offline augmentation and 458 duplicates were removed, splits were cut per capture session with an embargo, and perceptual hashing guarantees zero near-duplicates cross a split.
Solution
A FastAPI + WebSocket backend for real-time inference and a web frontend showing letters, text and text-to-speech; a full metrics report (precision/recall/F1, confusion matrix) can be regenerated at any time.
Model Details
- Problem
- Classifying 26 BISINDO fingerspelled letters from webcam frame sequences.
- Dataset
- BISINDO fingerspelling dataset, cleaned of offline augmentation and duplicates.
- Dataset Size
- 13,729 images · 26 classes
- Model
- MobileNetV3-Small + landmark MLP branch + Transformer encoder (2 layers, 4 heads)
- Validation Strategy
- 70/15/15 split per capture session with an 8-frame embargo and perceptual-hash (dHash) verification.
- Deployment
- FastAPI + WebSocket, in-browser webcam frontend.