From patchwork Mon Jul 12 05:09:26 2021 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Lingjiang Fang X-Patchwork-Id: 28907 Delivered-To: andriy.gelman@gmail.com Received: by 2002:a25:bbc9:0:0:0:0:0 with SMTP id c9csp2610505ybk; Sun, 11 Jul 2021 22:09:51 -0700 (PDT) X-Google-Smtp-Source: ABdhPJzjBLb6Rrk7BlQEpACOZIB4+5OvGh6BaHQ4qFN2yaxhFmyJfvElOEa4vFTDESHyfiqe8gF1 X-Received: by 2002:a17:906:15c2:: with SMTP id l2mr50632122ejd.348.1626066591332; Sun, 11 Jul 2021 22:09:51 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1626066591; cv=none; d=google.com; s=arc-20160816; b=zALPr+qFL51D0/6BHUK/LkPVDXQ5plAx86JyBU617t279vlLpQ6jP+R36P+AdrXBcL geoyUPtePrkuQaSJZuhvv5mxrs06Gnk9LiGRxOI9AYo62V8KLa2b7aKTzYNwS0MC4d4p TKiQAqgxVuUR/f/Y3y409TqdMjWrO+NBd2zBCWzYKIqNXNZHT79J3VnIOsbB6L8k2z6N aVZdIhsUfcdouUzd8lBf6elMdzTIHSJQ91sJcaaHZ7n9xKS4YqrHP5Q9/SxF2FnxB4rB v1bRf5x+VZfiXgTedSYva4um6UAtLILUS9dM9D7fwNNSz0uG7RFnDrT4wq0zQ9cx9LJc Defw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:mime-version:date:to:from:message-id :dkim-signature:delivered-to; bh=cYbNmO95qwwQa79PUTYhqE9gTyDpBlpHeWRdAg5NdvE=; b=ZJEKSoflf7Re/ofa/RpyWTRcoMFjsdTAlZSP0OyGrEz8lYxhxkjcPQ8qF32sW/MNXA WzHA21iENf7RyWYUR6lIMiAJtoZtLf5mhfNUTTPhfayC7ZZN6tspAXbm+wb4DmrksQ// /gMpPE5G4/ZDCAEB15sHaOT+N2lZc1GxPKfkzzyGL0R3TPMk943+8Keq8ycJ2RTBn/Qf nOO9UZ1KtpYzkeES20Fn8f8vGPzB9MBWyyBLcrGZkrUf5AAZ18C1fa6+M5Kkm3rlU2Cp RrcZOah+peO6neLVAdGnJko+KXsrGq7M75X/oZJxkayDIpj3dW2WOgOzmyCFDPFzZfFb RsEQ== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@foxmail.com header.s=s201512 header.b=TcqIoL14; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=foxmail.com Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id n8si16078032edo.29.2021.07.11.22.09.50; Sun, 11 Jul 2021 22:09:51 -0700 (PDT) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@foxmail.com header.s=s201512 header.b=TcqIoL14; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=foxmail.com Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id EBD1868A36E; Mon, 12 Jul 2021 08:09:48 +0300 (EEST) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from out203-205-221-233.mail.qq.com (out203-205-221-233.mail.qq.com [203.205.221.233]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id F0729689953 for ; Mon, 12 Jul 2021 08:09:41 +0300 (EEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=foxmail.com; s=s201512; t=1626066577; bh=MmDrju5RGSqS+/rYlkcbEo4EsImlhvbPVoF7vRnlNQk=; h=From:To:Cc:Subject:Date:Reply-To; b=TcqIoL14kY+fcGLpdxhHMelN4JoFAbJW/mEPAWOb8XqJMCuv2vNr3XHol7xBwX93D MwCgFxb4LcyCgCxgvp1TkMkh07AEjWyNQfA95eD/MSEduUjeXQ5RKLvAH5iB1pB+bB C6y9ohhwPjcCfSFJXi+M325OxjJmk/skg964vUOU= Received: from localhost.localdomain ([14.17.22.74]) by newxmesmtplogicsvrszc9.qq.com (NewEsmtp) with SMTP id 26413E0C; Mon, 12 Jul 2021 13:09:36 +0800 X-QQ-mid: xmsmtpt1626066576tgh88000d Message-ID: X-QQ-XMAILINFO: MRR5Jod2qmrXU5Ngr1ktNI46ZXk6gCfebSlrHoGbHfbrgLb+3urrHhNQK03Rq6 /aORik7zuS75e5s5hjFqrurrtdPuLc0HYgJ/l8JJ6qfBj50zGcM9CI6iS7SamFgK8LIkWAXO6AaV QracMhzjlcUQm5QHges0Ncqilgqtt37GRR9oiNbeWt1GmsDijCWTYA3nHYZXeFgflK/zUIv87yGr eM+HpAyDI1DuIVV+tRbmoraJ7g2fuaRFYN7F7C8pT1ow1nSY0EGJZeb7gIelo+jlKlFdekpA7Gzt Hv7aXaRsgbh3idhGx4gGoO2+YzurJbimz71Okgw99Q+KgnB3MSd8qvC/AECEQiiMt8xU9wDrIKKY 00POsmMSMgvvwbLvmMj/f044Kbpj84pYI67CltXE8TG6SM7IstuJ9i4hQrfvK18/3jWoK4AUmgIK HPEKjAQ/fobOJnVLtsPNF8JAP4W8RoBsJgD8mxL09Dm5SVHgQyj13s9Tv+ZV+J0lhAoaBYvg397w rU8bbk43VnPB7aT8GKHngRtxVR5P8e5RM/i0PyMoYlOR7MQfIyg9p5UO7jcZH/mXqjM9TjELJHtu HUul8IHi3jmQvgvzTlx9IX5715wGN4vjEghB5nksuulbAdD7dFf/IMLKhf4Iz6jpeWWupJ3yfOcy inTpaoXGepa2F6HSbQ8OMcL3biV7fb0YYbmuL9yn6Mn5TJvr44TgqlgCk/t77S0O3aaCbZw6L7qA Dbx+NFWeILjFRhF35iC/44mC8Xk7MwLhwQo2cA/KqjM/r3j6THuN8HIanUDk3ec3QVrDhQCpoo7+ phPOzketV490Tzydb+yBgqbF0hPvDa4AljE5zVGau1YZ6BZe69ktBj From: Lingjiang Fang To: ffmpeg-devel@ffmpeg.org Date: Mon, 12 Jul 2021 13:09:26 +0800 X-OQ-MSGID: <20210712050926.20929-1-vacingfang@foxmail.com> X-Mailer: git-send-email 2.29.2 MIME-Version: 1.0 Subject: [FFmpeg-devel] [PATCH V4] lavf/vf_ocr: add subregion support X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Lingjiang Fang Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: 3rdT/d9oRACb Content-Length: 4337 follow comments from Steven Liu --- doc/filters.texi | 8 ++++++++ libavfilter/vf_ocr.c | 45 +++++++++++++++++++++++++++++++++++++++++++- 2 files changed, 52 insertions(+), 1 deletion(-) diff --git a/doc/filters.texi b/doc/filters.texi index d991c06628..f41ba0ce46 100644 --- a/doc/filters.texi +++ b/doc/filters.texi @@ -15457,6 +15457,14 @@ Set character whitelist. @item blacklist Set character blacklist. + +@item x, y +Set top-left corner of the subregion, in pixels, default is (0,0). + +@item w, h +Set width and height of the subregion, in pixels, +default is the bottom-right part from given top-left corner. + @end table The filter exports recognized text as the frame metadata @code{lavfi.ocr.text}. diff --git a/libavfilter/vf_ocr.c b/libavfilter/vf_ocr.c index 6de474025a..55f04b6592 100644 --- a/libavfilter/vf_ocr.c +++ b/libavfilter/vf_ocr.c @@ -33,6 +33,8 @@ typedef struct OCRContext { char *language; char *whitelist; char *blacklist; + int x, y, x_in, y_in; + int w, h, w_in, h_in; TessBaseAPI *tess; } OCRContext; @@ -45,6 +47,10 @@ static const AVOption ocr_options[] = { { "language", "set language", OFFSET(language), AV_OPT_TYPE_STRING, {.str="eng"}, 0, 0, FLAGS }, { "whitelist", "set character whitelist", OFFSET(whitelist), AV_OPT_TYPE_STRING, {.str="0123456789abcdefghijklmnopqrstuvwxyzABCDEFGHIJKLMNOPQRSTUVWXYZ.:;,-+_!?\"'[]{}()<>|/\\=*&%$#@!~ "}, 0, 0, FLAGS }, { "blacklist", "set character blacklist", OFFSET(blacklist), AV_OPT_TYPE_STRING, {.str=""}, 0, 0, FLAGS }, + { "x", "top x of sub region", OFFSET(x), AV_OPT_TYPE_INT, {.i64=0}, 0, INT_MAX, FLAGS }, + { "y", "top y of sub region", OFFSET(y), AV_OPT_TYPE_INT, {.i64=0}, 0, INT_MAX, FLAGS }, + { "w", "width of sub region", OFFSET(w), AV_OPT_TYPE_INT, {.i64=0}, 0, INT_MAX, FLAGS }, + { "h", "height of sub region", OFFSET(h), AV_OPT_TYPE_INT, {.i64=0}, 0, INT_MAX, FLAGS }, { NULL } }; @@ -93,6 +99,41 @@ static int query_formats(AVFilterContext *ctx) return ff_set_common_formats(ctx, fmts_list); } +static void check_fix(int *x, int *y, int *w, int *h, int pic_w, int pic_h) +{ + // 0 <= x < pic_w + if (*x >= pic_w) + *x = 0; + // 0 <= y < pic_h + if (*y >= pic_h) + *y = 0; + + if (*w == 0 || *w + *x > pic_w) + *w = pic_w - *x; + if (*h == 0 || *h + *y > pic_h) + *h = pic_h - *y; +} + +static int config_input(AVFilterLink *inlink) +{ + AVFilterContext *ctx = inlink->dst; + OCRContext *s = ctx->priv; + + s->x_in = s->x; + s->y_in = s->y; + s->w_in = s->w; + s->h_in = s->h; + check_fix(&s->x_in, &s->y_in, &s->w_in, &s->h_in, inlink->w, inlink->h); + if ( s->x_in != s->x || s->y_in != s->y || + (s->w != 0 && s->w_in != s->w) || (s->h != 0 && s->h_in != s->h)) { + av_log(s, AV_LOG_WARNING, "config error, subregion changed to " + "x=%d, y=%d, w=%d, h=%d\n", + s->x_in, s->y_in, s->w_in, s->h_in); + } + + return 0; +} + static int filter_frame(AVFilterLink *inlink, AVFrame *in) { AVDictionary **metadata = &in->metadata; @@ -102,8 +143,9 @@ static int filter_frame(AVFilterLink *inlink, AVFrame *in) char *result; int *confs; + // TODO(vacing): support expression result = TessBaseAPIRect(s->tess, in->data[0], 1, - in->linesize[0], 0, 0, in->width, in->height); + in->linesize[0], s->x_in, s->y_in, s->w_in, s->h_in); confs = TessBaseAPIAllWordConfidences(s->tess); av_dict_set(metadata, "lavfi.ocr.text", result, 0); for (int i = 0; confs[i] != -1; i++) { @@ -134,6 +176,7 @@ static const AVFilterPad ocr_inputs[] = { .name = "default", .type = AVMEDIA_TYPE_VIDEO, .filter_frame = filter_frame, + .config_props = config_input, }, { NULL } };