Version 0.2.0
MobileTransformers 0.2.0
First public release. Export a Hugging Face model, pull it onto a phone, and chat, retrieve, classify, fine-tune and merge — entirely on the device. No server, no inference API, no data leaving the phone.
Highlights
Export (Python)
- One command turns a Hugging Face checkpoint into a device package: ONNX inference graph, PEFT-enabled training graph with an optimiser, tokenizer, optional embedding stage, and a manifest.
- PEFT methods: LoRA, LoRA-XS, and MARS — our own method, sharing adapter components across layers so trainable parameters grow with rank rather than depth.
- Decoder and encoder tasks: text generation, sequence classification, feature extraction.
Android SDK — mobiletransformers-android
-
MobileTransformers.fromPretrained(...)resolves, downloads, verifies and atomically installs. - Generation — streaming, chat template, KV cache, two selectable ONNX Runtime engines.
- Fine-tuning — a real training loop with a live loss curve, surviving backgrounding.
- Merging — folds the trained adapter into the inference weights, on device.
- Retrieval — chunking, embedding and vector search over your own documents.
- Classification — encoder packages scored per label.
- Tool calling — model output validated against the app's allowlist, bound to an Android intent. Nothing runs without consent.
- Federated adapter exchange — default-off and consent-gated.
Published
- Six model packages: https://huggingface.co/mobiletransformers
- Native build artifacts hosted, so a fresh clone provisions itself with
make fetch-native-deps.
Getting started
make doctor && make setup
make fetch-native-deps && make android-build