From patchwork Mon Jan 10 14:58:33 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 33178 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a6b:cd86:0:0:0:0:0 with SMTP id d128csp2798429iog; Mon, 10 Jan 2022 06:58:56 -0800 (PST) X-Google-Smtp-Source: ABdhPJzk7b0hOH0Tvgkwk/DYUCLJ9eYws17IHneJl8VQ+6KsAVZHPXGYMcLkpM7yFPQgMfw5JLYT X-Received: by 2002:a50:d74e:: with SMTP id i14mr20719edj.243.1641826736326; Mon, 10 Jan 2022 06:58:56 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1641826736; cv=none; d=google.com; s=arc-20160816; b=QLjC9RHJrD2g3N1x7QRvmZB64HDn92HUekwycgzzvGJ04dvt0TsRhFSwuJ57fDt3mg /3BI+PFyaWaXIQ5zjXcnyHF9B0F9FEahXnjWHGA1Vv63G2gPChkaRcpPli4bbSLoASEp ak7dtSgvAQU+N+H6Vg4UcQE+cqzRM1zN2T9cHFwtxgBBCNxi89d/8VPdDOC/EPn+haiA GzHnZRhWAb3jzIrOOl8B/uR7UFNJvTjxKfueUnwPWNzyVHxX0Tq9l0kMjw41/hH7RckI tOjEYDcj5GBADE11nkXIpNlP9U+ld8O1GdiUwkWE4oi8AYZP1FIpB2UsDcSKaKDxRno5 rMhA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:mime-version:message-id:date :dkim-signature:delivered-to; bh=C5Ubohg5bHj6kg0BEQP97rmJyilLdlnRsHZHv/LWYds=; b=aE2Xz1DZrFD+5EU406ghBcO1lrY8PlH28Oz/blNV67VuV0OjWD8cmAX3A4bZMduu1y /YNhpKcHEBGKem1Vm/T+0QkFOo30OwtIRX1OghHbh7CnCukG4CZJLI/r/Z3dVTKDEeOJ QX0D3oG08bkQfsrVh5tHZWUlJsivKOF6PRR84r/l84+WV8iRPWAMlne/rAlqIs1kD9HM hS47OgZqYuEyLROF4ZUBQrSOsAG8ONhrxQe2944uejE3pRSIR7ilIGaOYLSH2xYDjFhe LD8YaYOAIaT41eFRzeFkH3JkwIc843LuURvM2+eusDPo8zokmORy83NRTG2FkzYo4lCy Jw+Q== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b="Mrx//8Io"; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id lz22si3347096ejb.367.2022.01.10.06.58.55; Mon, 10 Jan 2022 06:58:56 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b="Mrx//8Io"; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 2EFA268ADF2; Mon, 10 Jan 2022 16:58:53 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-ed1-f74.google.com (mail-ed1-f74.google.com [209.85.208.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id D77F868A744 for ; Mon, 10 Jan 2022 16:58:46 +0200 (EET) Received: by mail-ed1-f74.google.com with SMTP id z8-20020a056402274800b003f8580bfb99so10353792edd.11 for ; Mon, 10 Jan 2022 06:58:46 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:message-id:mime-version:subject:from:to:cc; bh=F7EyLm06knB9HPqr7TbosorOd9emWq1Ry/5vPDN5t9E=; b=Mrx//8IobFl3I/kNT9ukZSbdfp8RFs5ZUnYHEp76yjgqeGA8S3dC6w+XT0fjQ+0MgF 2IT/obF3aChiBoKVchDstTU0iGAAkdtO+wq7BeLWspse7OEMDGEg3xbqQXuORFBwjCy5 dE1hw9jObUBi1NopSCnoNiTet+VAt5Qr/OaBy47xEwdq7Wy6R8WS+I7ndhjacTzuJzVw K9bcULldsUHwiZRty31PfWgS2xUx1A0+zhs105QJ0ifftL34zd1Z9BJIhrY19PaUlQq6 UbmMtXnq/Hbhj3JJTGdxWQSk+GrvM3qgECpSHn5yzXlQV8GrVA5zls1eVjBsmWcYlJ/T k3ag== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:message-id:mime-version:subject:from:to:cc; bh=F7EyLm06knB9HPqr7TbosorOd9emWq1Ry/5vPDN5t9E=; b=53IOckBW/bgh8HKp/O2iKkY1+EifL1tDLvct+Ok8xSIkrV7x0kBxFGUGw6MITeAWPf mU54VhHGr9PFOd2mfEMIrDZnc4grnMZR9vcjSWRKnE/i1er62L73gyAw827gRRdGYdyr K7cMKlgyaDgm4GtajZ41VlLInFLr+W0wt6cBusVmOuWgSgGAdXhnPPLXgplTIuGLtbXI h63kmYYg7nMMdpYdIpzaI7D5ze3ApPPA+f4C8K46iUH3EyIeh/HtQACQbS0hcmoAuvM/ 8c1dufIB63Gi2QeklzS0aU6TYPmJ7y2+yFtghCnIBsbcpWAbwVbeDvj36M+rPTrxHxU7 bD8g== X-Gm-Message-State: AOAM533J/hqANIke7+dXZJGTR/PZgRwTUQJphnDHS5FMfIADht+cnOxF rLFvOqP3weC98YUMjjUseai8CfsRXG+IW1fV4/n9WKRnxZnPnDZ8QyYRwXZbCQEMPF3Unsla9Dx k6I7h+AJIIKRHhkfats7FXrRRGb0es++NtW4kUBBmmaxmXffn60foeb6U9uGIndh68Z2NAX0= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:b61d:d4c4:5dce:c0af]) (user=alankelly job=sendgmr) by 2002:a17:906:cc50:: with SMTP id mm16mr85224ejb.515.1641826725977; Mon, 10 Jan 2022 06:58:45 -0800 (PST) Date: Mon, 10 Jan 2022 15:58:33 +0100 Message-Id: <20220110145836.3449558-1-alankelly@google.com> Mime-Version: 1.0 X-Mailer: git-send-email 2.34.1.575.g55b058a8bb-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 1/4] libswscale: Re-factor ff_shuffle_filter_coefficients. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: EKxLK2UWhILF Make the code more readable, follow the style guide and propagate memory allocation errors. --- libswscale/swscale_internal.h | 2 +- libswscale/utils.c | 68 ++++++++++++++++++++--------------- 2 files changed, 40 insertions(+), 30 deletions(-) diff --git a/libswscale/swscale_internal.h b/libswscale/swscale_internal.h index 3a78d95ba6..26d28d42e6 100644 --- a/libswscale/swscale_internal.h +++ b/libswscale/swscale_internal.h @@ -1144,5 +1144,5 @@ void ff_sws_slice_worker(void *priv, int jobnr, int threadnr, #define MAX_LINES_AHEAD 4 //shuffle filter and filterPos for hyScale and hcScale filters in avx2 -void ff_shuffle_filter_coefficients(SwsContext *c, int* filterPos, int filterSize, int16_t *filter, int dstW); +int ff_shuffle_filter_coefficients(SwsContext *c, int* filterPos, int filterSize, int16_t *filter, int dstW); #endif /* SWSCALE_SWSCALE_INTERNAL_H */ diff --git a/libswscale/utils.c b/libswscale/utils.c index c5ea8853d5..52f07e1661 100644 --- a/libswscale/utils.c +++ b/libswscale/utils.c @@ -278,39 +278,47 @@ static const FormatEntry format_entries[] = { [AV_PIX_FMT_P416LE] = { 1, 1 }, }; -void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, int filterSize, int16_t *filter, int dstW){ +int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, + int filterSize, int16_t *filter, + int dstW) +{ #if ARCH_X86_64 - int i, j, k, l; + int i = 0, j = 0, k = 0; int cpu_flags = av_get_cpu_flags(); + if (!filter || dstW % 16 != 0) return 0; if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { - if ((c->srcBpc == 8) && (c->dstBpc <= 14)){ - if (dstW % 16 == 0){ - if (filter != NULL){ - for (i = 0; i < dstW; i += 8){ - FFSWAP(int, filterPos[i + 2], filterPos[i+4]); - FFSWAP(int, filterPos[i + 3], filterPos[i+5]); - } - if (filterSize > 4){ - int16_t *tmp2 = av_malloc(dstW * filterSize * 2); - memcpy(tmp2, filter, dstW * filterSize * 2); - for (i = 0; i < dstW; i += 16){//pixel - for (k = 0; k < filterSize / 4; ++k){//fcoeff - for (j = 0; j < 16; ++j){//inner pixel - for (l = 0; l < 4; ++l){//coeff - int from = i * filterSize + j * filterSize + k * 4 + l; - int to = (i) * filterSize + j * 4 + l + k * 64; - filter[to] = tmp2[from]; - } - } - } - } - av_free(tmp2); - } - } - } + if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { + int16_t *filterCopy = NULL; + if (filterSize > 4) { + if (!FF_ALLOC_TYPED_ARRAY(filterCopy, dstW * filterSize)) + return AVERROR(ENOMEM); + memcpy(filterCopy, filter, dstW * filterSize * sizeof(int16_t)); + } + // Do not swap filterPos for pixels which won't be processed by + // the main loop. + for (i = 0; i + 8 <= dstW; i += 8) { + FFSWAP(int, filterPos[i + 2], filterPos[i + 4]); + FFSWAP(int, filterPos[i + 3], filterPos[i + 5]); + } + if (filterSize > 4) { + // 16 pixels are processed at a time. + for (i = 0; i + 16 <= dstW; i += 16) { + // 4 filter coeffs are processed at a time. + for (k = 0; k + 4 <= filterSize; k += 4) { + for (j = 0; j < 16; ++j) { + int from = (i + j) * filterSize + k; + int to = i * filterSize + j * 4 + k * 16; + memcpy(&filter[to], &filterCopy[from], 4 * sizeof(int16_t)); + } + } + } + } + if (filterCopy) + av_free(filterCopy); } } #endif + return 0; } int sws_isSupportedInput(enum AVPixelFormat pix_fmt) @@ -1836,7 +1844,8 @@ av_cold int sws_init_context(SwsContext *c, SwsFilter *srcFilter, get_local_pos(c, 0, 0, 0), get_local_pos(c, 0, 0, 0))) < 0) goto fail; - ff_shuffle_filter_coefficients(c, c->hLumFilterPos, c->hLumFilterSize, c->hLumFilter, dstW); + if ((ret = ff_shuffle_filter_coefficients(c, c->hLumFilterPos, c->hLumFilterSize, c->hLumFilter, dstW)) != 0) + goto nomem; if ((ret = initFilter(&c->hChrFilter, &c->hChrFilterPos, &c->hChrFilterSize, c->chrXInc, c->chrSrcW, c->chrDstW, filterAlign, 1 << 14, @@ -1846,7 +1855,8 @@ av_cold int sws_init_context(SwsContext *c, SwsFilter *srcFilter, get_local_pos(c, c->chrSrcHSubSample, c->src_h_chr_pos, 0), get_local_pos(c, c->chrDstHSubSample, c->dst_h_chr_pos, 0))) < 0) goto fail; - ff_shuffle_filter_coefficients(c, c->hChrFilterPos, c->hChrFilterSize, c->hChrFilter, c->chrDstW); + if ((ret = ff_shuffle_filter_coefficients(c, c->hChrFilterPos, c->hChrFilterSize, c->hChrFilter, c->chrDstW)) != 0) + goto nomem; } } // initialize horizontal stuff From patchwork Mon Jan 10 14:58:34 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 33179 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a6b:cd86:0:0:0:0:0 with SMTP id d128csp2798581iog; Mon, 10 Jan 2022 06:59:07 -0800 (PST) X-Google-Smtp-Source: ABdhPJxDqv33ohe9uVJ26RK24AcOfQavKp3PlFtdlnbKgCxVTKBRphzYbUUkG0zs3fpfMRI2wXJa X-Received: by 2002:a17:907:e86:: with SMTP id ho6mr130714ejc.208.1641826746858; Mon, 10 Jan 2022 06:59:06 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1641826746; cv=none; d=google.com; s=arc-20160816; b=i8HckgQvoRA6NgMk6Cg1hGxRK4P11Qs5+9qykNbRMOU1LnE5Sz3chi9jA1Nm2p3usK s53/Sh/Kh1D1SsnGbhFy/kUoC12wOdLaVmQg/PsUpv1j6HhZ5Fo68b6muN+aLXi/sSe5 TlsEJbvXobJPVqXivPUSRPKk/UdZsvD2evMTwKQZKz+tDpdUvgc+pzg8QNaCwkNaFYIc +aVnhWYlzqjBPJhCwC0mw6K53jVo0wSPpUklb6CEFP5RUVQ6sf4/L2WwL+yqz+v+xMdc mouJNpyE/S9ycDoLEndi9aMaB8hdqGyAtvnr9ogigiK7L2wREOBn/+2DDnH6TMpr6HSy YKXQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:references:mime-version :message-id:in-reply-to:date:dkim-signature:delivered-to; bh=x9LiPJaPpJvnQaGJzIZiMTCoHCMnO59ewkfz5OiPRic=; b=VZQqt2N5aNd5atACchBHbuinJg9APIMGtYHfc/9S7hdlyVkWSlet74Bx5h38Zx279e ITc80rrD264bB2RZBNKAI4Vo7n8BR+J8GPhN++MacXMyu5/aVi7uRxlcLTaQrMnh67tU ee4TO92csSDIB+m2RuAiZN+11jSIHXOSH8GIuqsmybCaaKiqKhHV6zLmw9XGhGDy34BQ MDWJ9DzPwEXjkt9KVX6OW1P1dO3Z15FnSMFJGQfUfSu6CSFKbLx97aZlis8nnIzk7tqG K52umnkytdrjC9eOb8Ix9xxUET/F16NMu1T+1og3Cw9Lvv0w9ZCeR+sxc2xYI/NMvtVM UaYQ== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=jq7e29rf; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id z10si4110034edc.589.2022.01.10.06.59.06; Mon, 10 Jan 2022 06:59:06 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=jq7e29rf; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 3280268ACF5; Mon, 10 Jan 2022 16:59:04 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-ed1-f74.google.com (mail-ed1-f74.google.com [209.85.208.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 64EFC68ACF5 for ; Mon, 10 Jan 2022 16:58:57 +0200 (EET) Received: by mail-ed1-f74.google.com with SMTP id q15-20020a056402518f00b003f87abf9c37so10328055edd.15 for ; Mon, 10 Jan 2022 06:58:57 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:in-reply-to:message-id:mime-version:references:subject:from:to :cc; bh=Wwj+Sc9WmUDQyhLfHao/MLUaiJbRA7EqbvhSw1acUSY=; b=jq7e29rfJHILMiGjhYSYYCjnYqqllHTRFDarvSxPNAj4UfDMx+CChh9HYruEvf17wb SKofJCQTMlctx6KweF0pcjJNTS1Ub6jfbzV4uFWBXaizteb/DJfsENKEYAzGkKN7FA/M bWFme7h+Prvz7NabkwoBVYu+UrE33DtNbG+astOToi9Nqd3hrDjKS/qUnX7dO1sZ05za uOGthaloDEqpKzA+nIgimuyV/DdfOibC8WKXM4UO9cmplzzl3jQ3YhLEYA7qFMBOVSAr fKz5ZBul8np5P2aULBQY6u4Sa12yuL/7shP9qXXA40poFwJpAegy2akAucpXCfYVVtXs pelA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:in-reply-to:message-id:mime-version :references:subject:from:to:cc; bh=Wwj+Sc9WmUDQyhLfHao/MLUaiJbRA7EqbvhSw1acUSY=; b=wXQwdkJeUwfDGbRT91b6EzCmc040OOOjG0PON6gzDbF3psEIZa3CpNDwGx5NtAzREa pD78sr7gNu/Tg9kItbUrerQ3QR1vpcKazAzDKZTCVJMqQeMoCTZdQlWDHVv0l07Tcgea Eb+KWckYRgODP1Aw1/pH4s9rW6MafUlsz4Rd6th8lcZdzAjOm7Gy3RdhwjFfEVIRdcGR Lwjpn1as0czStSiMIgnvUAB/j04lAFIvcTednPucvF9JVoS7JN3FvIWDof27GylHTEqS r9GNtZf/1T2LILH//MvBHgOAuDNc46CN7ZMBYjuJ1oEU64THCaQl6gA68RDTwiIyggwl 1t7A== X-Gm-Message-State: AOAM533D2ZkrcQmaF2sxfWh3AeNDnQmgGVfzlQSkh0N/h6t3UmVU9IP7 TZTA2LXL4fO6H94Vb9evnT2ECF6fSbMwwthhM/gw29Ff4+e4T+OZzLXhIRkX0oXSizQQZgEy7yh V5dSJuZC33vpzbJqqF2ZXy0DEKgTqsNK6Npb17UQ3PdM/H8SC3ceXcFQKJpsupuR1ITnIxMs= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:b61d:d4c4:5dce:c0af]) (user=alankelly job=sendgmr) by 2002:a05:6402:3551:: with SMTP id f17mr70743edd.64.1641826736686; Mon, 10 Jan 2022 06:58:56 -0800 (PST) Date: Mon, 10 Jan 2022 15:58:34 +0100 In-Reply-To: <20220110145836.3449558-1-alankelly@google.com> Message-Id: <20220110145836.3449558-2-alankelly@google.com> Mime-Version: 1.0 References: <20220110145836.3449558-1-alankelly@google.com> X-Mailer: git-send-email 2.34.1.575.g55b058a8bb-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 2/4] libswscale: Avx2 hscale can process any input of size which is a multiple of 4. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: fkfRK/opRtdj The main loop processes blocks of 16 pixels. The tail processes blocks of size 4. --- libswscale/x86/scale_avx2.asm | 48 +++++++++++++++++++++++++++++++++-- 1 file changed, 46 insertions(+), 2 deletions(-) diff --git a/libswscale/x86/scale_avx2.asm b/libswscale/x86/scale_avx2.asm index 20acdbd633..dc42abb100 100644 --- a/libswscale/x86/scale_avx2.asm +++ b/libswscale/x86/scale_avx2.asm @@ -53,6 +53,9 @@ cglobal hscale8to15_%1, 7, 9, 16, pos0, dst, w, srcmem, filter, fltpos, fltsize, mova m14, [four] shr fltsized, 2 %endif + cmp wq, 16 + jl .tail_loop + mov countq, 0x10 .loop: movu m1, [fltposq] movu m2, [fltposq+32] @@ -97,11 +100,52 @@ cglobal hscale8to15_%1, 7, 9, 16, pos0, dst, w, srcmem, filter, fltpos, fltsize, vpsrad m6, 7 vpackssdw m5, m5, m6 vpermd m5, m15, m5 - vmovdqu [dstq + countq * 2], m5 + vmovdqu [dstq], m5 + add dstq, 0x20 add fltposq, 0x40 add countq, 0x10 cmp countq, wq - jl .loop + jle .loop + + sub countq, 0x10 + cmp countq, wq + jge .end + +.tail_loop: + movu xm1, [fltposq] +%ifidn %1, X4 + pxor xm9, xm9 + pxor xm10, xm10 + xor innerq, innerq +.tail_innerloop: +%endif + vpcmpeqd xm13, xm13 + vpgatherdd xm3,[srcmemq + xm1], xm13 + vpunpcklbw xm5, xm3, xm0 + vpunpckhbw xm6, xm3, xm0 + vpmaddwd xm5, xm5, [filterq] + vpmaddwd xm6, xm6, [filterq + 16] + add filterq, 0x20 +%ifidn %1, X4 + paddd xm9, xm5 + paddd xm10, xm6 + paddd xm1, xm14 + add innerq, 1 + cmp innerq, fltsizeq + jl .tail_innerloop + vphaddd xm5, xm9, xm10 +%else + vphaddd xm5, xm5, xm6 +%endif + vpsrad xm5, 7 + vpackssdw xm5, xm5, xm5 + vmovq [dstq], xm5 + add dstq, 0x8 + add fltposq, 0x10 + add countq, 0x4 + cmp countq, wq + jl .tail_loop +.end: REP_RET %endmacro From patchwork Mon Jan 10 14:58:35 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 33180 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a6b:cd86:0:0:0:0:0 with SMTP id d128csp2798720iog; Mon, 10 Jan 2022 06:59:17 -0800 (PST) X-Google-Smtp-Source: ABdhPJx8yCfnqJPBN0D8iTBDO16x71g/sdFIFlROJP+emmp7MTEPHC690watl0yBqpYQ9zYxjvx+ X-Received: by 2002:a17:906:cc84:: with SMTP id oq4mr115525ejb.736.1641826757412; Mon, 10 Jan 2022 06:59:17 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1641826757; cv=none; d=google.com; s=arc-20160816; b=F0y3rbNy68kgTGLNNJ7KM2uqYUitLwBadvEomV62uZug2LYQrFJ4Gbrl/RIiPMrjQP T6YwYX2+ZsT4B7u1ON96H6QUu5t5o4zVMxeDCjrTvFHciOEjk/cKYMKRKu84J1OYudpr UFPBeKBspUzw+ZgwBwXbsRHnyxY8Xy+ZhNFSppm1SWjivnXf5COOaadWdxABvpp8JJe7 yrkdMrMNsKKpYtlesq5piMf1XJ+F8b+EJ6YK22niOe9Sa6UPyWltNtV3jxkTErMFZoJS l0PhTJ4sjJK4k+uln/t/fEOmfxVTe1UVlVqs7wE57246AHGyglG0G4QKXhJdUp6bCvjK os/A== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:references:mime-version :message-id:in-reply-to:date:dkim-signature:delivered-to; bh=P8kpr92oiDvG+/wIo5HifvHHFxkT5CXq2JkAUFLVr3I=; b=hiI2Qu38et2cxup34eRYtQFafi776snbHbsKu+hPfkREj8SWtYVly+JhJOr/fgKfa7 u0bX4QSn4tnkl2eRBuX53qqSqHF9yAwH+LLKkNY9kmjrqvl3rMymXeCblrTaBWC6MIaG OTkpj8c+hcl4rOOqkfoNv81nTmlWVsrWQSVHu5AWGejAHdMDNxK6Svb9B/45UtBf/HOT HMJmvOBVt3eKBoKb9oYLWt/EbZVkJDcyQIHKLQuUqxUvXNgrocX/j/esvtyWrOUKTWZg 3BowYcFVZTw+rs98y/Iy9nBF1lytySwkwH8cY/k4bX5vwogGxKUWuo5LMbLx/xVwrRVa TmzA== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=V25xPhZM; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id gm1si4072335ejc.26.2022.01.10.06.59.15; Mon, 10 Jan 2022 06:59:17 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=V25xPhZM; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 3E3BF68AECC; Mon, 10 Jan 2022 16:59:12 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-ed1-f74.google.com (mail-ed1-f74.google.com [209.85.208.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 5D4BD68AE5D for ; Mon, 10 Jan 2022 16:59:05 +0200 (EET) Received: by mail-ed1-f74.google.com with SMTP id eg24-20020a056402289800b003fe7f91df01so1337333edb.6 for ; Mon, 10 Jan 2022 06:59:05 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:in-reply-to:message-id:mime-version:references:subject:from:to :cc; bh=jrHugP6v7dsuKUxpxGBJoCEdFcEJ6JmdezHPsNp7Pjw=; b=V25xPhZMwvsrApb489niSJ8ZDNnC2bvdJXLe4JhAAtgKXS9tuYOYZrG77kxyS+exPA Qkvolr/pWp9VJMx9xljF3tfJGizPzVS6SPuqYYy5r5AIpd2eM7AHvkfuo7d2rhQemgBZ sa1Eelr6NBqtM400Iis9sfxsyjF8VRO4R/MGACuPnVqhEGFAg90nsydQGBQXHFy+D4v5 lAHOHdNLcDnW3YYpyxlp2vCb7CSUgC9LZ+66/pX5P49zyqu97ml5lHtLQ2pcbrG30B/9 7Nu7kOFLNJE6mv6O8Vq+p4xPyP2e0HbkYmKS8sgj5Ci+WWbQ5HP9egIann4Rfc5Od9Qd aBmg== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:in-reply-to:message-id:mime-version :references:subject:from:to:cc; bh=jrHugP6v7dsuKUxpxGBJoCEdFcEJ6JmdezHPsNp7Pjw=; b=Un9ioaASYnpEcIewzoaSm6DfE2LzmwFrlvDSI47IwC+xEZ9u6iY74SL1V+COSV9PRe GyQhVFRBCoJ5Ncn/m3p/VGy+gpSxN3q6SN/hnIq87KTjLXo+mzfBX+4Reqoanv97BfSP 2acRjjeYQVGiD0AtuQxEb5Sik91Ak9JXx/M/e7Y5ciONWv3S2l8yiglNsenUi0kSUlZF 1XbdBZCTtFKluTzP4yYC654ROCLVpHGRXdUkwnBov3iHcMnAUzFqa2K+ibL3Pr8IVoos n3Zf+FOm10jDYu5iehpeSjGkGyUfx2SDFuJ5sODsmuPkxYKm1Vds4cLnL7GUFRkyjQ5J D+tA== X-Gm-Message-State: AOAM5325rhaDQ/Nd+A6Refo6PmkNuhRfHCA8mTrVczp5XxyW/NrypyY/ QuuVXMf67qHjNicIbXntHjMhtBr3ad5f9Ty19lxKpFc+s14gLeX5FMEUwJMvCuna/L+J5mXIu8m u8EemZUQvKkC7rv/RQc7n3AnL7/wCxebywpQrask7j1lipv2mWWqEIv3yyf+QxVqe6pwTc4c= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:b61d:d4c4:5dce:c0af]) (user=alankelly job=sendgmr) by 2002:a05:6402:1908:: with SMTP id e8mr71936edz.22.1641826744861; Mon, 10 Jan 2022 06:59:04 -0800 (PST) Date: Mon, 10 Jan 2022 15:58:35 +0100 In-Reply-To: <20220110145836.3449558-1-alankelly@google.com> Message-Id: <20220110145836.3449558-3-alankelly@google.com> Mime-Version: 1.0 References: <20220110145836.3449558-1-alankelly@google.com> X-Mailer: git-send-email 2.34.1.575.g55b058a8bb-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 3/4] libswscale: Enable hscale_avx2 for input sizes which ar emultiples of 4. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: LbrbCsIjSHge ff_shuffle_filter_coefficients shuffles the tail as required. --- libswscale/utils.c | 17 +++++++++++++++-- libswscale/x86/swscale.c | 4 ++-- 2 files changed, 17 insertions(+), 4 deletions(-) diff --git a/libswscale/utils.c b/libswscale/utils.c index 52f07e1661..7e1e9c3834 100644 --- a/libswscale/utils.c +++ b/libswscale/utils.c @@ -285,7 +285,7 @@ int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, #if ARCH_X86_64 int i = 0, j = 0, k = 0; int cpu_flags = av_get_cpu_flags(); - if (!filter || dstW % 16 != 0) return 0; + if (!filter || (dstW % 4 != 0)) return 0; if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { int16_t *filterCopy = NULL; @@ -296,9 +296,11 @@ int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } // Do not swap filterPos for pixels which won't be processed by // the main loop. - for (i = 0; i + 8 <= dstW; i += 8) { + for (i = 0; i + 16 <= dstW; i += 16) { FFSWAP(int, filterPos[i + 2], filterPos[i + 4]); FFSWAP(int, filterPos[i + 3], filterPos[i + 5]); + FFSWAP(int, filterPos[i + 10], filterPos[i + 12]); + FFSWAP(int, filterPos[i + 11], filterPos[i + 13]); } if (filterSize > 4) { // 16 pixels are processed at a time. @@ -312,6 +314,17 @@ int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } } } + // 4 pixels are processed at a time in the tail. + for (; i + 4 <= dstW; i += 4) { + // 4 filter coeffs are processed at a time. + for (k = 0; k + 4 <= filterSize; k += 4) { + for (j = 0; j < 4; ++j) { + int from = (i + j) * filterSize + k; + int to = i * filterSize + j * 4 + k * 4; + memcpy(&filter[to], &filterCopy[from], 4 * sizeof(int16_t)); + } + } + } } if (filterCopy) av_free(filterCopy); diff --git a/libswscale/x86/swscale.c b/libswscale/x86/swscale.c index fdc93866a6..1d8f19aa5a 100644 --- a/libswscale/x86/swscale.c +++ b/libswscale/x86/swscale.c @@ -580,9 +580,9 @@ switch(c->dstBpc){ \ if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { - if (c->chrDstW % 16 == 0) + if (c->chrDstW % 4 == 0) ASSIGN_AVX2_SCALE_FUNC(c->hcScale, c->hChrFilterSize); - if (c->dstW % 16 == 0) + if (c->dstW % 4 == 0) ASSIGN_AVX2_SCALE_FUNC(c->hyScale, c->hLumFilterSize); } } From patchwork Mon Jan 10 14:58:36 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 33181 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a6b:cd86:0:0:0:0:0 with SMTP id d128csp2798851iog; Mon, 10 Jan 2022 06:59:27 -0800 (PST) X-Google-Smtp-Source: ABdhPJyqS73IsWfpMKzC1xiihn7xMDaE79CHH55DK3FZaunEzcdQXWQ5h55lijgzmHGUhIf1dZ8/ X-Received: by 2002:a05:6402:5107:: with SMTP id m7mr63891edd.108.1641826767075; Mon, 10 Jan 2022 06:59:27 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1641826767; cv=none; d=google.com; s=arc-20160816; b=lXUvMcLJG4X0T0N4i/0q1U3phYWvEyddMGvyKYOk2sV2hor3MVNsBZ1hOr5y480fof Fi4v55YrBnCSwFGigSs+JC9eoE72GF4q0ErCrmGIf00df1BFetdTs+9j3W+Nb8QJmAg+ Wt7z8OdSFO0yM45bIfINSpc90tfsaOYgQ+cefbcrX146uKbXQg0lxe1o8B/BN0AVZ1tc I1n/Pm7Lqg9OsgQsFRuQ1B+zqE+JKXeApPDNeOsV4oIBSsxT56Ym837X3nesAimFMAH2 YOn3pRn4ulTN+5uwVbfbos8XyC/y2w7rl+0HkhkjaYtV2AVwsxR+6XjimLUD01w5QyEM 2+QQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:references:mime-version :message-id:in-reply-to:date:dkim-signature:delivered-to; bh=YFVX2AJI7NMC82SqFkhTwy3jhUXpM8zYhqr9O8LI3sg=; b=Ggo2ydfgg1r6abfNg9X/1+gPId9oWFfWb7smtPIB4CVPS8xGR0aKaco3pZZFuGfANZ qX2leac0aOYe8HDthZEYjjicW09kCL530B9t5DYMnw2g+hFqTv8H2N42O6AjiLVmB/PS p2wGiJyE3xieTk9MRVw6k3E0aEEPS8nR5m1TUIjNn7K8WrGnsj2rTw0mMHasoZay79u5 /aO3q7jvlosZL4Vp/tQl+zr7Btbk0qMn8dXjqTmptoPkkY6X3gh075yQ3jx8ZOfh2xs+ vcKokvl8gllmlW5y8ALiBbJh5cywy0UzBtrKBWIsrURidwXMw4S3OMSiaGU0IMe03KKc lLzg== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=SNTvzphZ; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id gn8si3967301ejc.807.2022.01.10.06.59.26; Mon, 10 Jan 2022 06:59:27 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=SNTvzphZ; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 2D92768AEFF; Mon, 10 Jan 2022 16:59:20 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-yb1-f201.google.com (mail-yb1-f201.google.com [209.85.219.201]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id E92A468AEBD for ; Mon, 10 Jan 2022 16:59:13 +0200 (EET) Received: by mail-yb1-f201.google.com with SMTP id w35-20020a25ac23000000b006106b0711f2so19938579ybi.23 for ; Mon, 10 Jan 2022 06:59:13 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:in-reply-to:message-id:mime-version:references:subject:from:to :cc; bh=C+nJ4kP9XfLtJoVryrGXlyKmgISJGld81H/kTOL7jFk=; b=SNTvzphZgsMUmbJ9UAj/nPhfyljIo13npGzz9xJrFytAnM40sg++RWxujLEYHtyXdq Nq8X34nrvZ2plIwiTVJfQYrqUvjt6GNkqt9ZA3cktIZ29R0HcEDYoshxWITZ0133B75K ZMGUk4btK8lFwHHNpmroOiIv3UXd5OoqRA512yNdE4rM6h3fdLyGgAkMmrfxaUKSMssM 3o1B/qWJ9vLfoWlfYuJ+ume893ZvCPlsgxrSK5iJ9iOwBK/h8wGvK1IE76lRP4evcl3e epGMt9OsUoi3yvX3d34EMimUKfHkojJ8GY970lFhCAUPNd5iXRPIi+ehFtvH43UsPCmK sOOQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:in-reply-to:message-id:mime-version :references:subject:from:to:cc; bh=C+nJ4kP9XfLtJoVryrGXlyKmgISJGld81H/kTOL7jFk=; b=Wp7AGPDQLV3uo/yPWk730xrQAuWYKJUYUn1y0iFBCOii6mEWzO1BsqtgNmOdrjrCx1 gALvR1YCSKciMyJClVp8ZNYE5hctMillQeohIp/qkbO+CovGYB/+vqksHgzLdvHKyWJR b/4ieZoZAez13iqIu5bZorbTMp4z0KESf1gCel74JxOx/9rAFlmjZ4cEz959Xe5/PAVg EyQrhRorO0EeE6g5c0rJyUPfzY5Kr0w467+nwfA3baU0177vA1k2zqnk6oyLZTIhPDyJ /eJbz4oTo+Olb7enyurFQKCi/mjiefVxpGWX+00Iwr5PPTZvmEgYaesYqhVoYDtbovWK K0gw== X-Gm-Message-State: AOAM530EbHtVEx5zQBB+wc/RYsEJyX4NnsC3MMxBsyPv8nY0rDwV6NFp tsT/Ut+DTU9FGtdmFz83dE4LQv9b/+7HnTq0b6DR+r7Gs1j+ywKLFReSPPbgllBQg/Wu1CRI6QO U9RWg51X56HEuOdcoJq7h1pDf7yJERuOzoPPvLsg+XGFS65DR5fliDBymUvmWayUJN8wWCG8= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:b61d:d4c4:5dce:c0af]) (user=alankelly job=sendgmr) by 2002:a05:6902:1001:: with SMTP id w1mr86253742ybt.664.1641826752493; Mon, 10 Jan 2022 06:59:12 -0800 (PST) Date: Mon, 10 Jan 2022 15:58:36 +0100 In-Reply-To: <20220110145836.3449558-1-alankelly@google.com> Message-Id: <20220110145836.3449558-4-alankelly@google.com> Mime-Version: 1.0 References: <20220110145836.3449558-1-alankelly@google.com> X-Mailer: git-send-email 2.34.1.575.g55b058a8bb-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 4/4] checkasm/sw_scale: hscale does not requires cpuflag test. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: LROQE4BrsaRh This is done in ff_shuffle_filter_coefficients. --- tests/checkasm/sw_scale.c | 6 ++---- 1 file changed, 2 insertions(+), 4 deletions(-) diff --git a/tests/checkasm/sw_scale.c b/tests/checkasm/sw_scale.c index 3c0a083b42..e7f916d3a8 100644 --- a/tests/checkasm/sw_scale.c +++ b/tests/checkasm/sw_scale.c @@ -168,8 +168,6 @@ static void check_hscale(void) const uint8_t *src, const int16_t *filter, const int32_t *filterPos, int filterSize); - int cpu_flags = av_get_cpu_flags(); - ctx = sws_alloc_context(); if (sws_init_context(ctx, NULL, NULL) < 0) fail(); @@ -215,10 +213,10 @@ static void check_hscale(void) filter[SRC_PIXELS * width + i] = rnd(); } + ff_sws_init_scale(ctx); memcpy(filterAvx2, filter, sizeof(uint16_t) * (SRC_PIXELS * MAX_FILTER_WIDTH + MAX_FILTER_WIDTH)); - if ((cpu_flags & AV_CPU_FLAG_AVX2) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) - ff_shuffle_filter_coefficients(ctx, filterPosAvx, width, filterAvx2, SRC_PIXELS); + ff_shuffle_filter_coefficients(ctx, filterPosAvx, width, filterAvx2, SRC_PIXELS); if (check_func(ctx->hcScale, "hscale_%d_to_%d_width%d", ctx->srcBpc, ctx->dstBpc + 1, width)) { memset(dst0, 0, SRC_PIXELS * sizeof(dst0[0]));