Applied AI
On-device speech and language models, and the evaluation work that decides whether a feature ships.
Written response within one week. No preliminary sales call.
Typical engagement
- Typical duration
- 2–6 months
- Team
- 2–3 people
- Starts with
- Feasibility sprint
Indicative, not a quotation.
Engagement termsOverview.
The operational side of machine learning, which determines whether a feature holds up in use: evaluation, deployment and cost control, rather than model architecture.
An important part of this work is identifying when a model is not the right solution. A rules engine that can be debugged is often preferable to a model that cannot be explained, and the recommendation is made on that basis.
What the work covers.
The capabilities applied ai engagements draw on. Scope is agreed per engagement, against the requirement.
On-device speech recognition
Whisper-based transcription on Android and iOS, with model download, storage and lifecycle managed in the app.
On-device language models
Summaries, action items and question answering from a local model, with no upload.
Models sized to the device
Speech and text models tiered to each phone’s memory and processor, with features reduced gracefully on smaller devices.
Voice input surfaces
Dictation through a floating mic over any app or a dedicated keyboard.
Text clean-up on the device
Fixing, rewriting, changing tone and translating text in a field without a server round trip.
Guardrailed model decisions
Language models used behind deterministic rules and hard limits, with a written rationale recorded for each decision.
Feasibility and evaluation
An evaluation set and named failure classes agreed before build, and model choice justified against them.
What you receive.
What changes hands during the engagement, and remains yours after it.
Handed over
- Evaluation notes on model choice: size, latency and accuracy on target devices
- A model packaging and download strategy
- On-device or service integration in the product
- A guardrail and human-review path
Technologies in use.
Evidenced in our own products and tooling for this service. Where your team already runs a stack, the work proceeds within it.
- On-device inference
- whisper.cpp
- llama.cpp
- sherpa-onnx
- LiteRT
- Integration
- Kotlin Multiplatform
- Kotlin/Native C interop
- Llamatik
- Model formats
- GGUF
How it is delivered.
The sequence an engagement follows, from the first assessment onward.
Step 1: Feasibility sprint
Establish whether a model is the right tool, and define the evaluation set and failure classes first.
Step 2: Model and runtime selection
On the device or on a server, sized to the latency, privacy and cost budget.
Step 3: Prototype on target hardware
Measure the candidate models on the devices customers own.
Step 4: Production integration
Integrate with guardrails and a human-review path.
Step 5: Hold
Model updates and quality regression checks.
Evidence in our own work.
Products we engineered and operate under our own name, and tooling we publish. Client engagements ship under our clients’ names and are not listed.
Where this is not the right fit.
We do not add a conversational interface to a product that does not require one, and we do not deploy a model into a decision that materially affects a person’s life without a human review path.
Questions on applied ai.
Commercial, ownership and engagement questions are answered on the FAQ page.
All questionsDo you train models from scratch?
Rarely, and it is seldom advisable. Selecting and integrating an existing model is the appropriate approach in most cases.
How do you decide if it is good enough to ship?
Against a held-out evaluation set agreed with you before work begins, with the failure classes named in advance. Setting the threshold after seeing the results is how teams justify shipping systems that do not perform.
Can it run without sending data to a server?
Often, yes. MBS ships speech recognition and language-model features that run on the handset.
How do you handle large model downloads on mobile?
Models are downloaded and managed after install rather than bundled, with the user able to restrict downloads to Wi-Fi.
Will the model make decisions on its own?
Not where it matters. Deterministic rules and hard limits keep a veto, and each decision carries a rationale.
Other core services.
Mobile apps
Android and iOS apps, native or on a shared codebase where the product and the team justify one.
Games
Casual and children’s games, from playable prototype through store submission and live operations.
Web platforms
Dashboards, product sites and the data layers behind them, built to perform on a mid-range phone.
Desktop software
Windows, macOS and Linux applications for workflows that do not belong in a browser.
Design & planning
Product definition, scope and interface design, delivered as a fixed-fee engagement.
Something else?
Describe the requirement. Any other application or product is scoped against it.
Request a proposal.
Submit the requirement for applied ai work and receive a written assessment, including an indicative estimate and the risks identified at this stage.
Or write to hello@mobilebytesensei.com

