Building a Flutter document scanner requires an early architecture decision: run the OCR model on-device, or send images to a dedicated server. Both are valid. Neither is always right.
For Textify AI, a document intelligence application that extracts, summarizes, and translates content from physical documents, PixelPlot integrated PaddleOCR through a Python Flask backend. The application also uses BART for summarization and FLAN-T5 for question-and-answer generation.
This post covers how that server-side architecture is structured, and why on-device OCR is often impractical for production apps targeting a range of devices. For a broader look at what we build, see our AI and automation services.
On-device OCR using TensorFlow Lite or ONNX exports of PaddleOCR is technically possible. The tradeoffs make it difficult to recommend for most cases.
PaddleOCR's multilingual models are large. Bundling them as Flutter assets adds meaningful binary size. Inference on a mid-range Android device is noticeably slower than on server-grade hardware, particularly for dense or complex documents. On-device also means the model version is tied to the app release cycle.
Server-side OCR introduces a network dependency. If the server is unavailable, the feature is unavailable. For applications where internet access is assumed, this is an acceptable tradeoff. For field tools operating without connectivity, it is not.
The Flask service receives an image, runs PaddleOCR, and returns the extracted text with per-line confidence scores. Model initialization happens once at server startup, not per request. PaddleOCR loads several hundred megabytes of model weights. Reloading those on every request would make the endpoint unusable.
from flask import Flask, request, jsonify
from paddleocr import PaddleOCR
import base64
import numpy as np
import cv2
app = Flask(__name__)
# Initialize once outside the request handler
ocr = PaddleOCR(use_angle_cls=True, lang='en')
@app.route('/extract', methods=['POST'])
def extract_text():
data = request.get_json()
if not data or 'image' not in data:
return jsonify({'error': 'No image provided'}), 400
try:
img_bytes = base64.b64decode(data['image'])
np_arr = np.frombuffer(img_bytes, np.uint8)
img = cv2.imdecode(np_arr, cv2.IMREAD_COLOR)
result = ocr.ocr(img, cls=True)
lines = []
for block in result:
for line in block:
text, confidence = line[1]
lines.append({'text': text, 'confidence': round(confidence, 3)})
return jsonify({'lines': lines})
except Exception as e:
return jsonify({'error': str(e)}), 500use_angle_cls=True enables PaddleOCR's angle classification model, which detects rotated text before passing it to the recognition engine. This is useful for documents photographed at an angle but adds inference time. For applications where documents are consistently flat and horizontal, it can be disabled.
The image is compressed before encoding to avoid unnecessary payload size over mobile connections. A full-resolution JPEG from a modern smartphone camera is far larger than PaddleOCR requires for reliable text extraction.
import 'dart:convert';
import 'dart:io';
import 'package:http/http.dart' as http;
import 'package:image/image.dart' as img;
Future<List<OcrLine>> extractText(File imageFile, String serverUrl) async {
final rawBytes = await imageFile.readAsBytes();
final decoded = img.decodeImage(rawBytes);
if (decoded == null) throw Exception('Could not decode image');
// Resize to limit payload — PaddleOCR does not require higher resolution
final resized = img.copyResize(decoded, width: 1280);
final compressed = img.encodeJpg(resized, quality: 85);
final encoded = base64Encode(compressed);
final response = await http.post(
Uri.parse('$serverUrl/extract'),
headers: {'Content-Type': 'application/json'},
body: jsonEncode({'image': encoded}),
);
if (response.statusCode != 200) {
throw Exception('OCR error: ${response.statusCode}');
}
final responseData = jsonDecode(response.body) as Map<String, dynamic>;
final lines = (responseData['lines'] as List<dynamic>)
.cast<Map<String, dynamic>>()
.map((l) => OcrLine(
text: l['text'] as String,
confidence: (l['confidence'] as num).toDouble(),
))
.toList();
return lines;
}The target width of 1280px is a practical starting point, not a derived measurement. Adjust based on the document density and character size relevant to your use case.
PaddleOCR performs well on printed, high-contrast text. Its accuracy degrades on:
OpenCV's adaptive thresholding before passing the image to the OCR engine can improve results on low-contrast documents by normalizing pixel intensity. It will not reliably recover degraded handwriting.
The use_angle_cls model corrects moderate rotation well. Extreme perspective distortion (a document photographed at a steep angle from the side) requires a perspective transform correction step before OCR, which is a meaningful addition to the preprocessing pipeline.
Server-side OCR means document images leave the device. For applications handling sensitive documents, this requires clear disclosure in the privacy policy and a defined data retention policy for the server. Processing synchronously without persistence is a straightforward approach: the image is received, processed, and discarded within the same request. For higher compliance requirements, on-device inference or a self-hosted server under the client's data sovereignty is worth the architectural cost.
Textify AI also integrates Google Translate and TTS APIs for multilingual accessibility, which are additional services that handle user content on external infrastructure.
Technically yes, via TFLite or ONNX exports, but the models are large and inference time on mid-range Android hardware is substantially slower than on a server. A backend approach is more practical for most production applications where internet access is assumed.
Either approach works. Base64 JSON keeps the API interface uniform and simple. Multipart upload is more bandwidth-efficient for large payloads. The choice depends on whether bandwidth or implementation simplicity matters more for the specific deployment.
If the server scales to zero on inactivity, the first request after an idle period will be slow because Python must reload the PaddleOCR model into memory. A health-check ping on a short interval keeps the process warm. Alternatively, configure the server to maintain a minimum instance count.
On printed, high-contrast text in supported languages, PaddleOCR performs well. For handwritten text or degraded documents, accuracy drops and preprocessing becomes necessary. Evaluate against your actual document types before committing to it. For English-only use cases, Tesseract with LSTM mode is worth benchmarking as a comparison.