When the TexOps team approached us, their requirement sounded simple on paper: build a tablet app that textile workers can use to scan fabric and instantly detect manufacturing defects.
But there was a catch. Most of these textile mills operate in rural areas with spotty, unreliable internet. We couldn't just throw an image payload at a cloud endpoint like OpenAI or AWS Rekognition and wait for a JSON response. The computer vision model had to run entirely on-device, locally, without a network connection.
Here is a breakdown of how we architected the Flutter app to handle heavy C++ inference, the state management decisions we made to keep the UI fluid at 60fps, and the tradeoffs we accepted along the way.
We initially debated building this natively in Swift and Kotlin. Computer vision relies heavily on hardware acceleration (CoreML on iOS, NNAPI on Android).
However, we ultimately chose Flutter because we needed to rapidly iterate on the UI with the client's floor managers. The ability to push a single UI update across both tablet fleets (iPad and Samsung Galaxy Tabs) saved weeks of synchronization overhead.
But Flutter is a UI framework, not a machine learning engine. To bridge the gap, we bypassed standard platform channels (which introduce serialization overhead) and used Dart FFI (Foreign Function Interface) to talk directly to a C++ TensorFlow Lite wrapper.
If you use standard MethodChannel to pass a 4K image buffer from Dart to native code, you are going to drop frames. The serialization cost is just too high for 30fps video processing.
Instead, we allocated the camera buffer directly in shared memory using ffi:
// Allocating shared memory for the image buffer
final pointer = calloc<Uint8>(imageBytes.length);
final byteList = pointer.asTypedList(imageBytes.length);
byteList.setAll(0, imageBytes);
// Calling our C++ inference engine directly
final defectResult = _runInference(pointer, width, height);By passing only the memory pointer to the C++ engine, the UI thread never locks up, and memory duplication is entirely avoided.
Running a heavy neural network, even in C++, will still consume CPU cycles. If we ran this on the main thread, the entire Flutter UI would stutter every time a frame was processed.
We used Riverpod combined with Dart's Isolate.spawn to keep the UI perfectly decoupled from the math.
WorkerIsolate listens to this stream.(x, y, width, height, confidence).@riverpod
class DefectScanner extends _$DefectScanner {
@override
Stream<List<Defect>> build() async* {
final isolate = await ScannerIsolate.spawn();
await for (final frame in ref.watch(cameraFramesProvider)) {
yield await isolate.processFrame(frame);
}
}
}This separation meant that even if the inference took 150ms on an older tablet, the UI thread remained completely free to animate button presses, handle navigation, and keep the user experience buttery smooth.
You don't build an offline-first computer vision app without making sacrifices.
.apk and .ipa size by 45MB. We considered downloading the models on first launch, but the mills' internet made this too risky. We accepted the larger initial download.TexOps is now deployed across three major textile facilities. The app processes frames in roughly 42ms on an iPad Pro, instantly highlighting tears, stains, and weave defects without ever making a network request.
If you are dealing with a similar challenge - needing high-performance local processing without sacrificing a unified cross-platform UI - reach out to us. We'd love to talk architecture.