From patchwork Mon Jul 12 05:11:28 2021 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Lingjiang Fang X-Patchwork-Id: 28903 Delivered-To: andriy.gelman@gmail.com Received: by 2002:a25:bbc9:0:0:0:0:0 with SMTP id c9csp2611598ybk; Sun, 11 Jul 2021 22:11:45 -0700 (PDT) X-Google-Smtp-Source: ABdhPJzmZeH3BNsj6AUpM1XBqv0kOmKcxET0OvT7ZM+IVDtYNF2qT9SuOZwtA6UWLehb6k5FvaCS X-Received: by 2002:a17:907:3e22:: with SMTP id hp34mr4416071ejc.334.1626066705622; Sun, 11 Jul 2021 22:11:45 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1626066705; cv=none; d=google.com; s=arc-20160816; b=zn3Mu1k2pFxUNDDld/sw/fDxLCldvc/35mo0+KtfppDS7PhfXvuZJG2OkVHs5N715N slWNTV3a95ZZPK+4Bpl3uxOCsSP2UTHlqP4KerVcXqH8ZGs4CieTRnK98gEtJYISPcGf Nc9FUCW5/pFB5/cjfyrE3fwIL2HAveE6d0aaTDYGak9/j3waHWc7L93Cf4ZedZeN375e AFvOf3Izu2+6j0QS87lShJUkAdNCUMUE5NFUXoTf71p9bWANgguu4g11fK3v8kysg7sP paQO0+LPJnbhJ7/dMAH/370nKaPqI7ud5ZQBCk4f2cyHi4cXR020yt+313ARvIFE7pvF hVFQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:mime-version:date:to:from:message-id :dkim-signature:delivered-to; bh=cYbNmO95qwwQa79PUTYhqE9gTyDpBlpHeWRdAg5NdvE=; b=bUkh1y6b3apPcpopbXlu0gvmXDaakiRL0RXK1RwjV4mBM9BSrSQh8mtRi+gS4WmMia 94Jtv22p4+lo9d54YpyLtco+1w5EIJjHia3NzGGrM7Lf7NUeilZNdY4wKHxI996f0OHB XNhEu47GmdTzdDhvSHfycBaSgq3dNlzFME4/5dsRyd5BcsQlAF6XISnRFc9HOoTa0a/m JbB6uAdxRk/reXQVhipAJZMhw+uUmtNZP4yPeiuL3qBWslwqZdp48PMRH1VJDstsc0/P yxalNKMDuWXe3m9zHt+YXprObdwQM0rj29XuPTHMtsCUTBbApG+jJZ4TMWzaz3JNsDyp 1f2A== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@foxmail.com header.s=s201512 header.b=FO3atRCQ; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=foxmail.com Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id a2si13510706edb.549.2021.07.11.22.11.45; Sun, 11 Jul 2021 22:11:45 -0700 (PDT) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@foxmail.com header.s=s201512 header.b=FO3atRCQ; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=foxmail.com Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 4605268A591; Mon, 12 Jul 2021 08:11:44 +0300 (EEST) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from out203-205-221-173.mail.qq.com (out203-205-221-173.mail.qq.com [203.205.221.173]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id D88D468A0D7 for ; Mon, 12 Jul 2021 08:11:36 +0300 (EEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=foxmail.com; s=s201512; t=1626066691; bh=MmDrju5RGSqS+/rYlkcbEo4EsImlhvbPVoF7vRnlNQk=; h=From:To:Cc:Subject:Date:Reply-To; b=FO3atRCQ/b9D3XNHGi/rc6m/6nLrVeM9pgo7tgyhYHc0MjvUxt4Vxi8kwlNBqKYc2 2tRr5tH+Jj+OlTJ5sZP5DymWX7UgKrA8N/F1hxCyNQqOZlq30QT84QuwuNd5bIxDsm xrwQ3yhtGolG4ucaThWcu+Cw0kLfyPM2lUQDMpyA= Received: from localhost.localdomain ([14.17.22.74]) by newxmesmtplogicsvrsza28.qq.com (NewEsmtp) with SMTP id 2DE95458; Mon, 12 Jul 2021 13:11:30 +0800 X-QQ-mid: xmsmtpt1626066690td2l6b22o Message-ID: X-QQ-XMAILINFO: NqPEjZpmuGEFMhVPxU/96qbZftrbNJcsvDToPut9GjChAhY/UJxdEPzaMM3oDR qNzMRb6mxiPy/WIoanhp/IY0cjK47nWOcHb/TfgwPxrixan7ggqLcbjJBVxVdQJ6W8ZuMWkjtqmD Vnjx6joxVWyOhzVVnxBGorWPl62iuG8XZNSdaLA4nChq4EK0XuVuxUiPgfUZmPHT/XPvtuCIfS/m 6Ood0ty3zkE7lP0No6A2aiZ6NG5JPyy5XPQPe6LXReWma/ShXTlys/JvHNBGMy4RAj4Z1TpFxdX6 Lwt3aIq33QARoI5vNN3ZHI9QX1sSsgDTW0WHLIfpINI4LBrBVXW/yI5xeO0ZKXpXE7xw7ABzUAlj E9UJ0S0Iw8Xxs0+ZzBvqYunmz31wNBfQe+KfRCEdpG9TG7MbXoIP+mHKUSiaRynIJzuSmNCXYcrZ fN2EdgPsP2nfMYSoykpy7NscvyXHrzCAf8odCDgylLMKkefL51ytKpKtuJTtIv4CbO13Uokzgnr2 E99vu1BiK+dtosrTlQOf2a8E2aI++8TzAS2BvF/3juaPuw7ZhZNtfzC0aeuQYookGXB0jfgNQ29H TgKQ9BncllB3SV5l0qoQEhS7SsXixmBgI87PaCr8IjBUA9qAvjMua9OzosPUiqs5S+apFvfROY7P Oppt6bn41ykdvd0T1sIs3cfifGpelcEhBzGkjX22VVLSxcKFm1Eo4M9d0rXdrOOK9XXVDIBSFLQV wEZQh3GUIo/SaF5vuPTn9QUY+vI6gIUQkwDpc/vQEc5hie90ksbhCFTKempGT10FSly/ohmOEluG s5stbWMsEDvb2nvGqYOLIIUDACPhnuNhVEHGEMT8r/uxnGwAEAO/nuKFjRRQTWPOQQDXb4ufAdIE nad6UByjjRyvy/+u7JmX4= From: Lingjiang Fang To: ffmpeg-devel@ffmpeg.org Date: Mon, 12 Jul 2021 13:11:28 +0800 X-OQ-MSGID: <20210712051128.23021-1-vacingfang@foxmail.com> X-Mailer: git-send-email 2.29.2 MIME-Version: 1.0 Subject: [FFmpeg-devel] [PATCH V4] lavf/vf_ocr: add subregion support X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Lingjiang Fang Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: xsCZpljpTOba Content-Length: 4337 follow comments from Steven Liu --- doc/filters.texi | 8 ++++++++ libavfilter/vf_ocr.c | 45 +++++++++++++++++++++++++++++++++++++++++++- 2 files changed, 52 insertions(+), 1 deletion(-) diff --git a/doc/filters.texi b/doc/filters.texi index d991c06628..f41ba0ce46 100644 --- a/doc/filters.texi +++ b/doc/filters.texi @@ -15457,6 +15457,14 @@ Set character whitelist. @item blacklist Set character blacklist. + +@item x, y +Set top-left corner of the subregion, in pixels, default is (0,0). + +@item w, h +Set width and height of the subregion, in pixels, +default is the bottom-right part from given top-left corner. + @end table The filter exports recognized text as the frame metadata @code{lavfi.ocr.text}. diff --git a/libavfilter/vf_ocr.c b/libavfilter/vf_ocr.c index 6de474025a..55f04b6592 100644 --- a/libavfilter/vf_ocr.c +++ b/libavfilter/vf_ocr.c @@ -33,6 +33,8 @@ typedef struct OCRContext { char *language; char *whitelist; char *blacklist; + int x, y, x_in, y_in; + int w, h, w_in, h_in; TessBaseAPI *tess; } OCRContext; @@ -45,6 +47,10 @@ static const AVOption ocr_options[] = { { "language", "set language", OFFSET(language), AV_OPT_TYPE_STRING, {.str="eng"}, 0, 0, FLAGS }, { "whitelist", "set character whitelist", OFFSET(whitelist), AV_OPT_TYPE_STRING, {.str="0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ.:;,-+_!?\"'[]{}()<>|/\\=*&%$#@!~ "}, 0, 0, FLAGS }, { "blacklist", "set character blacklist", OFFSET(blacklist), AV_OPT_TYPE_STRING, {.str=""}, 0, 0, FLAGS }, + { "x", "top x of sub region", OFFSET(x), AV_OPT_TYPE_INT, {.i64=0}, 0, INT_MAX, FLAGS }, + { "y", "top y of sub region", OFFSET(y), AV_OPT_TYPE_INT, {.i64=0}, 0, INT_MAX, FLAGS }, + { "w", "width of sub region", OFFSET(w), AV_OPT_TYPE_INT, {.i64=0}, 0, INT_MAX, FLAGS }, + { "h", "height of sub region", OFFSET(h), AV_OPT_TYPE_INT, {.i64=0}, 0, INT_MAX, FLAGS }, { NULL } }; @@ -93,6 +99,41 @@ static int query_formats(AVFilterContext *ctx) return ff_set_common_formats(ctx, fmts_list); } +static void check_fix(int *x, int *y, int *w, int *h, int pic_w, int pic_h) +{ + // 0 <= x < pic_w + if (*x >= pic_w) + *x = 0; + // 0 <= y < pic_h + if (*y >= pic_h) + *y = 0; + + if (*w == 0 || *w + *x > pic_w) + *w = pic_w - *x; + if (*h == 0 || *h + *y > pic_h) + *h = pic_h - *y; +} + +static int config_input(AVFilterLink *inlink) +{ + AVFilterContext *ctx = inlink->dst; + OCRContext *s = ctx->priv; + + s->x_in = s->x; + s->y_in = s->y; + s->w_in = s->w; + s->h_in = s->h; + check_fix(&s->x_in, &s->y_in, &s->w_in, &s->h_in, inlink->w, inlink->h); + if ( s->x_in != s->x || s->y_in != s->y || + (s->w != 0 && s->w_in != s->w) || (s->h != 0 && s->h_in != s->h)) { + av_log(s, AV_LOG_WARNING, "config error, subregion changed to " + "x=%d, y=%d, w=%d, h=%d\n", + s->x_in, s->y_in, s->w_in, s->h_in); + } + + return 0; +} + static int filter_frame(AVFilterLink *inlink, AVFrame *in) { AVDictionary **metadata = &in->metadata; @@ -102,8 +143,9 @@ static int filter_frame(AVFilterLink *inlink, AVFrame *in) char *result; int *confs; + // TODO(vacing): support expression result = TessBaseAPIRect(s->tess, in->data[0], 1, - in->linesize[0], 0, 0, in->width, in->height); + in->linesize[0], s->x_in, s->y_in, s->w_in, s->h_in); confs = TessBaseAPIAllWordConfidences(s->tess); av_dict_set(metadata, "lavfi.ocr.text", result, 0); for (int i = 0; confs[i] != -1; i++) { @@ -134,6 +176,7 @@ static const AVFilterPad ocr_inputs[] = { .name = "default", .type = AVMEDIA_TYPE_VIDEO, .filter_frame = filter_frame, + .config_props = config_input, }, { NULL } };