Running YOLOv8 on-device inside a Flutter app is not complicated in theory. In practice, there are a few places where things go wrong, and they are not well-documented. This is what we encountered building the defect detection module for TexOps, a Flutter-based textile mill ERP.
The requirement was real-time fabric defect detection that worked without an internet connection. Most textile facilities in the region have unreliable connectivity, and cloud inference was off the table from day one.
The model itself was a YOLOv8n (nano) variant fine-tuned on fabric defect images. Training happened in Python with Ultralytics. Getting it into TFLite is straightforward with the export command:
from ultralytics import YOLO
model = YOLO("texops_defect_v2.pt")
model.export(format="tflite", int8=True, data="fabric_defects.yaml")The int8=True flag is important for mobile. It reduces the model size from about 12MB to just under 4MB, and inference time on a mid-range Android tablet dropped from roughly 280ms per frame to about 85ms. That is the difference between a detection that feels sluggish and one that feels real-time during a fabric inspection.
The quantization did reduce accuracy slightly. On our validation set, the full float32 model hit 91.3% mAP and the int8 version hit 88.7%. For identifying obvious manufacturing defects (tears, stains, missed weave patterns) at the coarse level the client needed, 88.7% was acceptable. If the application had been medical imaging, it would not have been.
We used the tflite_flutter package. The tricky part is that the package's documentation shows how to run inference but glosses over memory management for continuous video processing.
The first version of our inference code created a new Interpreter object on every camera frame. On a newer device this was fine. On the older Galaxy Tab A tablets some floor workers were using, it caused memory pressure that eventually crashed the app after a few minutes of scanning.
The fix was to create the interpreter once at app startup and hold a reference to it in a singleton service:
class DefectInferenceService {
static DefectInferenceService? _instance;
late Interpreter _interpreter;
static Future<DefectInferenceService> getInstance() async {
_instance ??= await _init();
return _instance!;
}
static Future<DefectInferenceService> _init() async {
final service = DefectInferenceService();
final modelData = await rootBundle.load('assets/models/texops_defect_v2.tflite');
service._interpreter = Interpreter.fromBuffer(modelData.buffer.asUint8List());
return service;
}
List<Defect> runInference(CameraImage frame) {
// Reuse the same interpreter instance rather than creating a new one per frame
final input = _preprocessFrame(frame);
final output = List.filled(1 * 25200 * 8, 0.0).reshape([1, 25200, 8]);
_interpreter.run(input, output);
return _parseOutput(output);
}
}After this change, the app ran for extended shifts without memory issues on the older tablets.
YOLOv8 expects a 640x640 normalized float32 input. Camera frames from Flutter's camera plugin come in as YUV420 format. Converting them correctly is where we lost the most time.
The obvious approach was to use Flutter's image package to decode the YUV and resize. This worked, but it was slow enough that we were only processing about 4-5 frames per second on slower hardware, which felt unresponsive during scanning.
We ended up writing the YUV-to-RGB conversion and resizing as a Dart isolate operation to keep it off the main thread, and caching the normalization constants to avoid recalculating them per frame. Getting from YUV to a clean normalized tensor is genuinely unglamorous work, and the documentation for doing it correctly in Flutter is scattered.
The defect detection screen shows a live camera view with bounding boxes drawn over detected defects. Lab engineers scan a moving section of fabric by passing the tablet slowly over it. Defects above the confidence threshold (set to 0.6 in the production build) trigger an audio cue and get logged automatically to Firestore with a timestamp and a thumbnail crop.
Inference runs at about 12 frames per second on the older tablets and closer to 20 on the newer iPads. Both are fast enough for the inspection workflow. The model is bundled in the app assets, so it works immediately with no download step and no connectivity requirement.
If you are building on-device vision features in Flutter, the most important lessons from our side: do your model export with quantization enabled, hold a single interpreter instance for the app lifecycle, and do not trust that frame preprocessing will be fast by default. Those three things account for most of the rough edges we hit.