Skip to content

Version 0.2.0

Martin Korelič requested to merge lrk into main

MobileTransformers 0.2.0

First public release. Export a Hugging Face model, pull it onto a phone, and chat, retrieve, classify, fine-tune and merge — entirely on the device. No server, no inference API, no data leaving the phone.

Highlights

Export (Python)

  • One command turns a Hugging Face checkpoint into a device package: ONNX inference graph, PEFT-enabled training graph with an optimiser, tokenizer, optional embedding stage, and a manifest.
  • PEFT methods: LoRA, LoRA-XS, and MARS — our own method, sharing adapter components across layers so trainable parameters grow with rank rather than depth.
  • Decoder and encoder tasks: text generation, sequence classification, feature extraction.

Android SDK — mobiletransformers-android

  • MobileTransformers.fromPretrained(...) resolves, downloads, verifies and atomically installs.
  • Generation — streaming, chat template, KV cache, two selectable ONNX Runtime engines.
  • Fine-tuning — a real training loop with a live loss curve, surviving backgrounding.
  • Merging — folds the trained adapter into the inference weights, on device.
  • Retrieval — chunking, embedding and vector search over your own documents.
  • Classification — encoder packages scored per label.
  • Tool calling — model output validated against the app's allowlist, bound to an Android intent. Nothing runs without consent.
  • Federated adapter exchange — default-off and consent-gated.

Published

Getting started

make doctor && make setup
make fetch-native-deps && make android-build

Docs: https://martinkorelic.github.io/mobiletransformers/

Merge request reports

Loading