From patchwork Mon Jan 10 14:58:35 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 33180 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a6b:cd86:0:0:0:0:0 with SMTP id d128csp2798720iog; Mon, 10 Jan 2022 06:59:17 -0800 (PST) X-Google-Smtp-Source: ABdhPJx8yCfnqJPBN0D8iTBDO16x71g/sdFIFlROJP+emmp7MTEPHC690watl0yBqpYQ9zYxjvx+ X-Received: by 2002:a17:906:cc84:: with SMTP id oq4mr115525ejb.736.1641826757412; Mon, 10 Jan 2022 06:59:17 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1641826757; cv=none; d=google.com; s=arc-20160816; b=F0y3rbNy68kgTGLNNJ7KM2uqYUitLwBadvEomV62uZug2LYQrFJ4Gbrl/RIiPMrjQP T6YwYX2+ZsT4B7u1ON96H6QUu5t5o4zVMxeDCjrTvFHciOEjk/cKYMKRKu84J1OYudpr UFPBeKBspUzw+ZgwBwXbsRHnyxY8Xy+ZhNFSppm1SWjivnXf5COOaadWdxABvpp8JJe7 yrkdMrMNsKKpYtlesq5piMf1XJ+F8b+EJ6YK22niOe9Sa6UPyWltNtV3jxkTErMFZoJS l0PhTJ4sjJK4k+uln/t/fEOmfxVTe1UVlVqs7wE57246AHGyglG0G4QKXhJdUp6bCvjK os/A== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:references:mime-version :message-id:in-reply-to:date:dkim-signature:delivered-to; bh=P8kpr92oiDvG+/wIo5HifvHHFxkT5CXq2JkAUFLVr3I=; b=hiI2Qu38et2cxup34eRYtQFafi776snbHbsKu+hPfkREj8SWtYVly+JhJOr/fgKfa7 u0bX4QSn4tnkl2eRBuX53qqSqHF9yAwH+LLKkNY9kmjrqvl3rMymXeCblrTaBWC6MIaG OTkpj8c+hcl4rOOqkfoNv81nTmlWVsrWQSVHu5AWGejAHdMDNxK6Svb9B/45UtBf/HOT HMJmvOBVt3eKBoKb9oYLWt/EbZVkJDcyQIHKLQuUqxUvXNgrocX/j/esvtyWrOUKTWZg 3BowYcFVZTw+rs98y/Iy9nBF1lytySwkwH8cY/k4bX5vwogGxKUWuo5LMbLx/xVwrRVa TmzA== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=V25xPhZM; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id gm1si4072335ejc.26.2022.01.10.06.59.15; Mon, 10 Jan 2022 06:59:17 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=V25xPhZM; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 3E3BF68AECC; Mon, 10 Jan 2022 16:59:12 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-ed1-f74.google.com (mail-ed1-f74.google.com [209.85.208.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 5D4BD68AE5D for ; Mon, 10 Jan 2022 16:59:05 +0200 (EET) Received: by mail-ed1-f74.google.com with SMTP id eg24-20020a056402289800b003fe7f91df01so1337333edb.6 for ; Mon, 10 Jan 2022 06:59:05 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:in-reply-to:message-id:mime-version:references:subject:from:to :cc; bh=jrHugP6v7dsuKUxpxGBJoCEdFcEJ6JmdezHPsNp7Pjw=; b=V25xPhZMwvsrApb489niSJ8ZDNnC2bvdJXLe4JhAAtgKXS9tuYOYZrG77kxyS+exPA Qkvolr/pWp9VJMx9xljF3tfJGizPzVS6SPuqYYy5r5AIpd2eM7AHvkfuo7d2rhQemgBZ sa1Eelr6NBqtM400Iis9sfxsyjF8VRO4R/MGACuPnVqhEGFAg90nsydQGBQXHFy+D4v5 lAHOHdNLcDnW3YYpyxlp2vCb7CSUgC9LZ+66/pX5P49zyqu97ml5lHtLQ2pcbrG30B/9 7Nu7kOFLNJE6mv6O8Vq+p4xPyP2e0HbkYmKS8sgj5Ci+WWbQ5HP9egIann4Rfc5Od9Qd aBmg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:in-reply-to:message-id:mime-version :references:subject:from:to:cc; bh=jrHugP6v7dsuKUxpxGBJoCEdFcEJ6JmdezHPsNp7Pjw=; b=Un9ioaASYnpEcIewzoaSm6DfE2LzmwFrlvDSI47IwC+xEZ9u6iY74SL1V+COSV9PRe GyQhVFRBCoJ5Ncn/m3p/VGy+gpSxN3q6SN/hnIq87KTjLXo+mzfBX+4Reqoanv97BfSP 2acRjjeYQVGiD0AtuQxEb5Sik91Ak9JXx/M/e7Y5ciONWv3S2l8yiglNsenUi0kSUlZF 1XbdBZCTtFKluTzP4yYC654ROCLVpHGRXdUkwnBov3iHcMnAUzFqa2K+ibL3Pr8IVoos n3Zf+FOm10jDYu5iehpeSjGkGyUfx2SDFuJ5sODsmuPkxYKm1Vds4cLnL7GUFRkyjQ5J D+tA== X-Gm-Message-State: AOAM5325rhaDQ/Nd+A6Refo6PmkNuhRfHCA8mTrVczp5XxyW/NrypyY/ QuuVXMf67qHjNicIbXntHjMhtBr3ad5f9Ty19lxKpFc+s14gLeX5FMEUwJMvCuna/L+J5mXIu8m u8EemZUQvKkC7rv/RQc7n3AnL7/wCxebywpQrask7j1lipv2mWWqEIv3yyf+QxVqe6pwTc4c= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:b61d:d4c4:5dce:c0af]) (user=alankelly job=sendgmr) by 2002:a05:6402:1908:: with SMTP id e8mr71936edz.22.1641826744861; Mon, 10 Jan 2022 06:59:04 -0800 (PST) Date: Mon, 10 Jan 2022 15:58:35 +0100 In-Reply-To: <20220110145836.3449558-1-alankelly@google.com> Message-Id: <20220110145836.3449558-3-alankelly@google.com> Mime-Version: 1.0 References: <20220110145836.3449558-1-alankelly@google.com> X-Mailer: git-send-email 2.34.1.575.g55b058a8bb-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 3/4] libswscale: Enable hscale_avx2 for input sizes which ar emultiples of 4. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: LbrbCsIjSHge ff_shuffle_filter_coefficients shuffles the tail as required. --- libswscale/utils.c | 17 +++++++++++++++-- libswscale/x86/swscale.c | 4 ++-- 2 files changed, 17 insertions(+), 4 deletions(-) diff --git a/libswscale/utils.c b/libswscale/utils.c index 52f07e1661..7e1e9c3834 100644 --- a/libswscale/utils.c +++ b/libswscale/utils.c @@ -285,7 +285,7 @@ int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, #if ARCH_X86_64 int i = 0, j = 0, k = 0; int cpu_flags = av_get_cpu_flags(); - if (!filter || dstW % 16 != 0) return 0; + if (!filter || (dstW % 4 != 0)) return 0; if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { int16_t *filterCopy = NULL; @@ -296,9 +296,11 @@ int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } // Do not swap filterPos for pixels which won't be processed by // the main loop. - for (i = 0; i + 8 <= dstW; i += 8) { + for (i = 0; i + 16 <= dstW; i += 16) { FFSWAP(int, filterPos[i + 2], filterPos[i + 4]); FFSWAP(int, filterPos[i + 3], filterPos[i + 5]); + FFSWAP(int, filterPos[i + 10], filterPos[i + 12]); + FFSWAP(int, filterPos[i + 11], filterPos[i + 13]); } if (filterSize > 4) { // 16 pixels are processed at a time. @@ -312,6 +314,17 @@ int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } } } + // 4 pixels are processed at a time in the tail. + for (; i + 4 <= dstW; i += 4) { + // 4 filter coeffs are processed at a time. + for (k = 0; k + 4 <= filterSize; k += 4) { + for (j = 0; j < 4; ++j) { + int from = (i + j) * filterSize + k; + int to = i * filterSize + j * 4 + k * 4; + memcpy(&filter[to], &filterCopy[from], 4 * sizeof(int16_t)); + } + } + } } if (filterCopy) av_free(filterCopy); diff --git a/libswscale/x86/swscale.c b/libswscale/x86/swscale.c index fdc93866a6..1d8f19aa5a 100644 --- a/libswscale/x86/swscale.c +++ b/libswscale/x86/swscale.c @@ -580,9 +580,9 @@ switch(c->dstBpc){ \ if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { - if (c->chrDstW % 16 == 0) + if (c->chrDstW % 4 == 0) ASSIGN_AVX2_SCALE_FUNC(c->hcScale, c->hChrFilterSize); - if (c->dstW % 16 == 0) + if (c->dstW % 4 == 0) ASSIGN_AVX2_SCALE_FUNC(c->hyScale, c->hLumFilterSize); } }