From patchwork Thu Feb 17 10:04:20 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 34358 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6838:d078:0:0:0:0 with SMTP id x24csp490252nkx; Thu, 17 Feb 2022 02:04:34 -0800 (PST) X-Google-Smtp-Source: ABdhPJx4fuNhfdqCR4ENcD/CxqnXMV+TdPxH/E2CEWw/LDwLt/e+bi4BKywyiidkghwMJzSmZWp+ X-Received: by 2002:a17:907:1a48:b0:6ce:4aa:30c5 with SMTP id mf8-20020a1709071a4800b006ce04aa30c5mr1701162ejc.559.1645092274209; Thu, 17 Feb 2022 02:04:34 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1645092274; cv=none; d=google.com; s=arc-20160816; b=M7SBNJFr4HkrrTSXcssU9ZfP3sqK82xp6xDdIlkG04rf3A3G7g5yyjQ1Ki1ZMmBQVj QkCfPKyP2PrSv/5WxJ+oTfB/GE6m1aoCVyj4jHbB/uytI1VAksNcKSkzDdfwo3IHUAYi dQV4hxNnudOuPqOqzz1vrVmdj4Fvh6cemwiP7WbzX0VtlnTLB46zM1XOSdxGcB2hZkto cdfT4LV+Tfv+sx5/85y5dHmJ8jEYzyemtTWdnwc95RlF3yb7Q786p8vTQ1QCVV0oaRl6 ZWHaGYibTqSyBB3cFtScfPs+Ydzf78/h1nuZu7xAGNv22xbG+Ll/TEM6/e0SdA/yJ9s2 Ughw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:mime-version:message-id:date :dkim-signature:delivered-to; bh=8pOlWr8dVQtm2nlYsE0eKY32m/hvY38QR8mbFDnMZ24=; b=qeBmXkDSJFae+XY78pGROGPjuhv7O1TAK/HbvSm2aQ07uwJTDxrStf7lJMXRbBZZ1/ b6biG8f53Q3C7oxS3SHtfJfXZtViv4IJZv9F2F8zLXp+7lBV6CQCK2scW90cOWwA2e72 4Yv9ALWdnVmOAD4VA9Dwbki1FL8P4jyMmP16L7uaC77P+ifQp+YB/nrLX8NM+WFysqt4 yezas3EcDXKbrZKujfLm3HQTptfTkRiR3F3iHcdRQ2a5lxkJ3IZNCZWZRA7VtrkLZH/P i9St+gqwCOsKl/gR7whhGeyKOiKyBdp6FTlj/k7dIAU/LQeaDd5+FRGj9hCwwSjL8Awv uDKg== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=PFfSyZbG; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id c12si1283446eja.863.2022.02.17.02.04.33; Thu, 17 Feb 2022 02:04:34 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=PFfSyZbG; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 1D0A168B1CC; Thu, 17 Feb 2022 12:04:30 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-ej1-f74.google.com (mail-ej1-f74.google.com [209.85.218.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id AAE5668B1CC for ; Thu, 17 Feb 2022 12:04:23 +0200 (EET) Received: by mail-ej1-f74.google.com with SMTP id d7-20020a1709061f4700b006bbf73a7becso1267927ejk.17 for ; Thu, 17 Feb 2022 02:04:23 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:message-id:mime-version:subject:from:to:cc; bh=sP+FgB7BKGXQp8fa1HxGM7u8B33V09CCnI8tayWyPN0=; b=PFfSyZbGD86CmFmIb55OEp0/75OvXGmRWswQGK83UgLGthPNuaEmyInOFGK/3G48B1 1h5YkeHKDyJ10vnbGqh3q5Xnp2/NGHVr+H7BuwTKVqhnXnEJ7ZzIsaTjdNy3GOsiWdWb 9uMlfq1w2afAMZY+SD44i8NG9eKtU7TmiJDlDyIgcaDeBuMRHLJQWgIZbB0WyjYdJx2D iPWdGJGeeB9/xIecgwJFFf3sMDzffzjxDqoA0XgSvh3qO1kAhT491C3VZkAl2akVS8T+ NlDbSwJjVs+GfP4XcRywcoNBxnXwaOP2JC3QVSuxVqkABneewpGwr5+wwrEnx/1dEeMN GvVg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:message-id:mime-version:subject:from:to:cc; bh=sP+FgB7BKGXQp8fa1HxGM7u8B33V09CCnI8tayWyPN0=; b=UMH0dfo1LRcb4Iu8h5c6ihEaLE8Lzg85DoaXWk/LXE/yTAtu3t3dJAKJUIeNG8ckiJ +ayKMwaUPIxXwh4+blVV8Y1e0AlvF58WN8V7B1jzLAcGV5MCLpCh/ODht+pRuIu2e0X8 MdhcJL8OxJeCQmPGcBnOu+wxmQwwqR58F1wwYdDClfJUq/yEDZzFfQYMtDaQG9ef7z8T UdjWNaS6AHR2i0/NIYckZk6zt2I6/9hlzopEzwEbxSmuvcaPkhQDaVxRFkB1wBnBTTcl ejEAf+ucv3VufMJ5dLAeBIMQxbEDTI+ngCLi4zDL1cDfgTAkfXXW4UI5pz0EPDsLQnIT jCOw== X-Gm-Message-State: AOAM533gVQaBVGFfLGb/1mkPSS8GWyaPnZfpixH7ZQ4eFV9OFi4pIWtz gkM/iWuIE3e0M7mRK0LSu6CisQ5UDGfVANCwPlphPdHz728ZvLkFCMOv1rSLJJYI7vT32xzm015 p0xNF87ad5syYECUIP/GZfEz6LR0cY3HF8UDvUGTW7Ax2xqtLl6WS8Z68NgCQqTRztv+MVOE= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:b159:808d:943e:13ba]) (user=alankelly job=sendgmr) by 2002:a17:906:743:b0:6d0:7f19:d737 with SMTP id z3-20020a170906074300b006d07f19d737mr557864ejb.11.1645092263098; Thu, 17 Feb 2022 02:04:23 -0800 (PST) Date: Thu, 17 Feb 2022 11:04:20 +0100 Message-Id: <20220217100420.1113388-1-alankelly@google.com> Mime-Version: 1.0 X-Mailer: git-send-email 2.35.1.265.g69c8d7142f-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH v2 4/5] libswscale: Enable hscale_avx2 for all input sizes. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: thqdtjSaKPvp ff_shuffle_filter_coefficients shuffles the tail as required. --- libswscale/utils.c | 19 ++++++++++++++++--- libswscale/x86/swscale.c | 6 ++---- 2 files changed, 18 insertions(+), 7 deletions(-) diff --git a/libswscale/utils.c b/libswscale/utils.c index 7c8e1bbdde..d818c9ce55 100644 --- a/libswscale/utils.c +++ b/libswscale/utils.c @@ -285,8 +285,7 @@ int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, #if ARCH_X86_64 int i, j, k; int cpu_flags = av_get_cpu_flags(); - // avx2 hscale filter processes 16 pixel blocks. - if (!filter || dstW % 16 != 0) + if (!filter) return 0; if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { @@ -298,9 +297,11 @@ int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } // Do not swap filterPos for pixels which won't be processed by // the main loop. - for (i = 0; i + 8 <= dstW; i += 8) { + for (i = 0; i + 16 <= dstW; i += 16) { FFSWAP(int, filterPos[i + 2], filterPos[i + 4]); FFSWAP(int, filterPos[i + 3], filterPos[i + 5]); + FFSWAP(int, filterPos[i + 10], filterPos[i + 12]); + FFSWAP(int, filterPos[i + 11], filterPos[i + 13]); } if (filterSize > 4) { // 16 pixels are processed at a time. @@ -314,6 +315,18 @@ int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } } } + // 4 pixels are processed at a time in the tail. + for (; i < dstW; i += 4) { + // 4 filter coeffs are processed at a time. + int rem = dstW - i >= 4 ? 4 : dstW - i; + for (k = 0; k + 4 <= filterSize; k += 4) { + for (j = 0; j < rem; ++j) { + int from = (i + j) * filterSize + k; + int to = i * filterSize + j * 4 + k * 4; + memcpy(&filter[to], &filterCopy[from], 4 * sizeof(int16_t)); + } + } + } } av_free(filterCopy); } diff --git a/libswscale/x86/swscale.c b/libswscale/x86/swscale.c index 73869355b8..76f5a70fc5 100644 --- a/libswscale/x86/swscale.c +++ b/libswscale/x86/swscale.c @@ -691,10 +691,8 @@ switch(c->dstBpc){ \ if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { - if (c->chrDstW % 16 == 0) - ASSIGN_AVX2_SCALE_FUNC(c->hcScale, c->hChrFilterSize); - if (c->dstW % 16 == 0) - ASSIGN_AVX2_SCALE_FUNC(c->hyScale, c->hLumFilterSize); + ASSIGN_AVX2_SCALE_FUNC(c->hcScale, c->hChrFilterSize); + ASSIGN_AVX2_SCALE_FUNC(c->hyScale, c->hLumFilterSize); } }