Halo
RM.
Computer Vision

JEMARI — Real-Time BISINDO Sign Translator

Translates Indonesian Sign Language (BISINDO) fingerspelling from a webcam into text and speech in real time with a CNN + Transformer.

Role
ML Engineer (research implementation)
Client
Felix Aria & Muhammad Raysar Al-Fatih — SMAN 1 Balikpapan (OPSI 2026)
Year
2026
Technologies
PythonPyTorchMediaPipeFastAPIJavaScript
JE

Overview

A working implementation of an OPSI 2026 research proposal: the full pipeline from dataset preprocessing and GPU training to a real-time web app.

The Problem

Digital communication remains hard to access for BISINDO signers; an automatic translator running directly from a camera was needed.

Approach

Hands are detected and cropped with MediaPipe; MobileNetV3-Small extracts per-frame features, fused with 21 hand-joint landmarks, and a Transformer encoder models 8 consecutive frames before classifying into 26 letters.

Challenges

The first version reported 96.99% accuracy, which turned out to be invalid: 56.9% of the dataset was offline augmentation, so the test set contained copies of training images. Offline augmentation and 458 duplicates were removed, splits were cut per capture session with an embargo, and perceptual hashing guarantees zero near-duplicates cross a split.

Solution

A FastAPI + WebSocket backend for real-time inference and a web frontend showing letters, text and text-to-speech; a full metrics report (precision/recall/F1, confusion matrix) can be regenerated at any time.

Model Details

Problem
Classifying 26 BISINDO fingerspelled letters from webcam frame sequences.
Dataset
BISINDO fingerspelling dataset, cleaned of offline augmentation and duplicates.
Dataset Size
13,729 images · 26 classes
Model
MobileNetV3-Small + landmark MLP branch + Transformer encoder (2 layers, 4 heads)
Validation Strategy
70/15/15 split per capture session with an 8-frame embargo and perceptual-hash (dHash) verification.
Deployment
FastAPI + WebSocket, in-browser webcam frontend.