From patchwork Wed Feb 9 09:09:45 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 34205 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6602:2c4e:0:0:0:0 with SMTP id x14csp1560273iov; Wed, 9 Feb 2022 01:10:06 -0800 (PST) X-Google-Smtp-Source: ABdhPJxfSZImUUnPE0zAI7wbtn1kbNtJiukQYDZZYXYEQXq9gmWsIzfsFS6N9UaTrEl8neToLKhT X-Received: by 2002:a17:907:6d83:: with SMTP id sb3mr1046213ejc.21.1644397805930; Wed, 09 Feb 2022 01:10:05 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1644397805; cv=none; d=google.com; s=arc-20160816; b=rkuwH+6bpdNAcIooi1Tm6rCL1aiZHT7RTxzpVAJgALqiv4OLTu6Y3+jB9kv7v+GWdo suWwFZqhHeLGgEo5TL+NaqFxO+dR2U5eJMF2CbtW7Uu/uvWK9MPEbOFQT0QsaKjVpK4J 3teLNxZL4HCFXi60Yzj06hQddNE+c1LZ+Pr7RJMvSovY3hiHw0ZWRKYJCSoSRouZoOvG T/V6ikKq1Vlc0SrjjWEYuu+I997lDSKLTnpKCQQOTmRTBk6f9gRW7jtuaU7Op+SGOx0Q ouK8gfQSMYq/785esBeMUV9W0q84qUZG/0trTqCX1pF7FyBvxyluROwOZ8Z9jHF/+k3V JmVQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:mime-version:message-id:date :dkim-signature:delivered-to; bh=Il8e0o0+6IOafK8+VSH+XvD5QR2BCX2kfpofxE2nbgc=; b=lhbMJPlH8Kvt/5HaCR5a0tdZR0ZQjLy6nDUM8Q1kD4C8wCaKW8rlBNyZJ4LvrPoVrT TQnyQtk2jQQguQWJ3p2uKtchLog7F9OVPe9O83ed2PLgHsxebduGaX11a/zDL4X403Ad 4hRszlHgPULD4b5NwAaDX9mh/tgtAndZILqmO9+niTSWg+LIFBIJ9sPC/qaO+X8UemeD NHTFoXYI5WphYLNWCqoH29QXRTnLStHwBqzhGqm/mnAx5/NqXJUPAx5ETXQKFgn4g+/Y rASoU96kfkIg3nG1zcfsGfDEaevtm1/TnTTc3vfGBn1972YoRNZxTT0p+35G4iOhvpmR AxTw== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=MMiLHsG5; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id 10si9627587ejc.890.2022.02.09.01.10.05; Wed, 09 Feb 2022 01:10:05 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=MMiLHsG5; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id AA59268B192; Wed, 9 Feb 2022 11:10:01 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-wm1-f74.google.com (mail-wm1-f74.google.com [209.85.128.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 7F35F68B05A for ; Wed, 9 Feb 2022 11:09:54 +0200 (EET) Received: by mail-wm1-f74.google.com with SMTP id l4-20020a05600c4f0400b0037bb2ce79d8so2339586wmq.9 for ; Wed, 09 Feb 2022 01:09:54 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:message-id:mime-version:subject:from:to:cc; bh=1MobYHjS9erwzffYOOiKPv85FxvpX2eCy3TXODXNNGU=; b=MMiLHsG5TUXaR/Oib/8wyUdMFttnGkUB5Zq1D6jV1pslHqKwmnEPuCjRdMxZAdUH9Y hy6HmtjwfQa4zCTlCMt00a9RjZNhjV0fB2oYsqxYkZpl7dVEzWkaPJ1cXRdhTiCg82+j r2LBfmrEC/8cg/AudvJ9/uM1/m6/FxbUgU6OpvyOPtG0xhCz3y2RXt7Y/QbUolWlfNfL Mwlbn2tWVXVoXaLd7HM/FtOkNeTXUSyIHBqBYYrqSHNgiFIDmYxG202pkY3H06ok6Fmy hGVvTVHSAdRUwTcVa2TINrOUORwmO4BHqCG6/XnpC5ybXtOgHGAM0AQY18IAPf9Q4V5h LxCw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:message-id:mime-version:subject:from:to:cc; bh=1MobYHjS9erwzffYOOiKPv85FxvpX2eCy3TXODXNNGU=; b=mnx47aawJBYCFe/gbevuujSitaEE5rLz8yqeUz2TliR+KBWJxV24OT7yF22JrEC08X ko5s6C0vOe7mwv5qyHByoeMQ/xCJgbvRVHtX+fEH66dUOZ8ez0mDUZWD6NTcpyy+mnlW h6iBTnZGD97ThDCfBBxvePDuPU791n/Wk7eauHsd4Pv+ObAPTOw3Ef7+EXcfhw7BUEDq hH666iWKvJCqC5wx4Qc1Vyjmo7abq634OIYdH0GrRZ+EjTUoVc5ka1z0vz7P6mBtTKKP TI+sstCDZkd1xSSncPA6X6qUQzhG1GE0j39Dm0cATJfAhFL53MpATHj6WW+3DNnkstbm eqEw== X-Gm-Message-State: AOAM533PoE609NKM1Wh5GgHJ+uNnZecQ+hqjIYegzXOGUbnZDjr8ZwZJ i2raB6A/b+ypOfK+RyEltLKMEnQ7mMk8ImiAHe/0D3pqXocfVo+a2fkBX4aYXARjJKIAnwZ3HIY pRhyuCl/jZaPJ2uSDZ9g6Aal38f1WaUYkVVpcDEPlXF65ihGNoNFYqxFsFlEus8V+bOkJbZg= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:388b:9d0c:1bc4:40c]) (user=alankelly job=sendgmr) by 2002:adf:f08b:: with SMTP id n11mr1246124wro.7.1644397793528; Wed, 09 Feb 2022 01:09:53 -0800 (PST) Date: Wed, 9 Feb 2022 10:09:45 +0100 Message-Id: <20220209090945.3450752-1-alankelly@google.com> Mime-Version: 1.0 X-Mailer: git-send-email 2.35.0.263.gb82422642f-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 1/5] libswscale: Re-factor ff_shuffle_filter_coefficients. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: CkTwsPUXy9ac Make the code more readable and follow the style guide. --- libswscale/utils.c | 64 +++++++++++++++++++++++++++------------------- 1 file changed, 37 insertions(+), 27 deletions(-) diff --git a/libswscale/utils.c b/libswscale/utils.c index c5ea8853d5..1d919e863a 100644 --- a/libswscale/utils.c +++ b/libswscale/utils.c @@ -278,39 +278,49 @@ static const FormatEntry format_entries[] = { [AV_PIX_FMT_P416LE] = { 1, 1 }, }; -void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, int filterSize, int16_t *filter, int dstW){ +void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, + int filterSize, int16_t *filter, + int dstW) +{ #if ARCH_X86_64 - int i, j, k, l; + int i, j, k; int cpu_flags = av_get_cpu_flags(); + // avx2 hscale filter processes 16 pixel blocks. + if (!filter || dstW % 16 != 0) + return; if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { - if ((c->srcBpc == 8) && (c->dstBpc <= 14)){ - if (dstW % 16 == 0){ - if (filter != NULL){ - for (i = 0; i < dstW; i += 8){ - FFSWAP(int, filterPos[i + 2], filterPos[i+4]); - FFSWAP(int, filterPos[i + 3], filterPos[i+5]); - } - if (filterSize > 4){ - int16_t *tmp2 = av_malloc(dstW * filterSize * 2); - memcpy(tmp2, filter, dstW * filterSize * 2); - for (i = 0; i < dstW; i += 16){//pixel - for (k = 0; k < filterSize / 4; ++k){//fcoeff - for (j = 0; j < 16; ++j){//inner pixel - for (l = 0; l < 4; ++l){//coeff - int from = i * filterSize + j * filterSize + k * 4 + l; - int to = (i) * filterSize + j * 4 + l + k * 64; - filter[to] = tmp2[from]; - } - } - } - } - av_free(tmp2); - } - } - } + if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { + int16_t *filterCopy = NULL; + if (filterSize > 4) { + if (!FF_ALLOC_TYPED_ARRAY(filterCopy, dstW * filterSize)) + return; + memcpy(filterCopy, filter, dstW * filterSize * sizeof(int16_t)); + } + // Do not swap filterPos for pixels which won't be processed by + // the main loop. + for (i = 0; i + 8 <= dstW; i += 8) { + FFSWAP(int, filterPos[i + 2], filterPos[i + 4]); + FFSWAP(int, filterPos[i + 3], filterPos[i + 5]); + } + if (filterSize > 4) { + // 16 pixels are processed at a time. + for (i = 0; i + 16 <= dstW; i += 16) { + // 4 filter coeffs are processed at a time. + for (k = 0; k + 4 <= filterSize; k += 4) { + for (j = 0; j < 16; ++j) { + int from = (i + j) * filterSize + k; + int to = i * filterSize + j * 4 + k * 16; + memcpy(&filter[to], &filterCopy[from], 4 * sizeof(int16_t)); + } + } + } + } + if (filterCopy) + av_free(filterCopy); } } #endif + return; } int sws_isSupportedInput(enum AVPixelFormat pix_fmt) From patchwork Wed Feb 9 09:13:51 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 34206 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6602:2c4e:0:0:0:0 with SMTP id x14csp1563024iov; Wed, 9 Feb 2022 01:14:06 -0800 (PST) X-Google-Smtp-Source: ABdhPJycZHbxm/JZlGwdxiU+j4sGXUrjbXAuGOcoEyOid29HMaB+sfaO+MuYzbeagWxTV9TgsvpD X-Received: by 2002:a17:907:6298:: with SMTP id nd24mr1080996ejc.76.1644398046122; Wed, 09 Feb 2022 01:14:06 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1644398046; cv=none; d=google.com; s=arc-20160816; b=qPGL8mhGasm8H1pUQrrSScWAFdA9A8RvNwJQg2/q2/omvIv0zByaSHJ0nC3ndtmBT8 NOtCsMeluwlDyUBqUHYYTS1QLR8AXUkMdjsKB/Lju0Ba0LjjBavJxxp0hTeugkc0GjkC zIdbwMKRnmRYxQbVTemqSzHENbAXVXTSMVXaYBij3Zj3XPjTXTK7fNvbaRswpBSxyR11 VT8UcAqS4UeTjlQqmXMdF+k2Deu0akLtyBhL9qtm3PVIwLehY9ymCmM2y10N/Mc1GnK2 O1eIGZJD7tRTww6jP23/m12RtFdJgz/VdURPLjfcVE8Gp2BeaHCnheZfGZ/5oTpT7NL9 lrJg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:mime-version:message-id:date :dkim-signature:delivered-to; bh=7tSzi7cbY6MB2nD6+AoB1ymAuLxcs6O8cMXpzKJ7s4I=; b=0aLXSNeUJ/2FL5A0lUG93Unz1qYmcer4KBiBQhXeXXkOEQLTjz2Kf4kqCzZG5jey7u vX8G2cJs6fBJb0K9I9JX8NvS1xpX9s0hMcqponkRNeiDXJwaYu5exxMk/DzHLbzeoBRx h0o+0pk3BJX0yb5T5Qt7jpgCu9uitNIMbAq2FonwC8NCzaoSrP0FaByufqXXxbsHFkOk PPfaBXT5hPonhDSrdWp+gTldjUNtnG7QbgRLOE9q2MDvVwdRRF685qthSY+msEwkmoyv 1cysztTtdQYLc8QwNV2hPr/Y14jrYH0bJSmhBhPlm6MLjvlLMc1ebBklC9nXtLv080vy aRCQ== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=IQfxLWkI; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id g5si1769498edy.112.2022.02.09.01.14.05; Wed, 09 Feb 2022 01:14:06 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=IQfxLWkI; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 087A668B1C0; Wed, 9 Feb 2022 11:14:03 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-wm1-f73.google.com (mail-wm1-f73.google.com [209.85.128.73]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 26FC168A888 for ; Wed, 9 Feb 2022 11:13:56 +0200 (EET) Received: by mail-wm1-f73.google.com with SMTP id r8-20020a7bc088000000b0037bbf779d26so226501wmh.7 for ; Wed, 09 Feb 2022 01:13:56 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:message-id:mime-version:subject:from:to:cc; bh=7RTLgeKUdGbFt2Vmrv/Q9Gw1Gqgkn2ZPPhlqZiOosec=; b=IQfxLWkIqK+jEnvMGT6fpRm79PUAytZZL3SQTgjL19qUJKzS0rR37kEzBcFxQkbWV7 PEkCkH13Dk1ZiGv12iVm0EXnsVlLWMwgpe+IFRlU5bXe3tsKIWkKm+zFbOqbZLQbQShR vXTNfFGPNuHVMCuVYu4ukmj6SXNmXODMZH0tjG8ZuLFGQn0/eGK3SzMBXUSUKMlcgQbK YnvvwaI1TCs4jjW9e2luU1ZrOux4ZbLg8SrldzqF0AD1OXCuIoQNaZkzlgX3I7fGfE3Y qJxYrBoqLDkIPGLXrhxnB0WZIEexVVUPaF+OrbpgvIOkvZCCTTSfjcZdOmUm16XjRhQ8 zh4w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:message-id:mime-version:subject:from:to:cc; bh=7RTLgeKUdGbFt2Vmrv/Q9Gw1Gqgkn2ZPPhlqZiOosec=; b=VxJi+n8mS1GcHlL9+aqvMq2Gbo7nTuHyyU2n87Pn1sMvPfb0b9t55fOoAWehaYTLkf zH8gCXMNNxjJoOUcZhCc/4z+vlKnCYdBNXKY5HJ06fwJHFz6zGrgQaGtnX1hzx/N41Kb MneVhHDmM/hTYKtTERpg598trN6zM9AvgXF3XQ5qoWYjc6nPEGaBKwY1w/dy7OlKc35N 26o+ZAPL+2X9wjGDEaKW7Sbur8gUPG5l7KIBgtu/sAw4VCGLa4iy1frcmRLfUEdiZYYX cqj3GbfBs9W+++1cXe+leBgXhieki0MYgQBYcJHF23Djy0NDxpWVU5qq+33c08u6z1Sq SwGA== X-Gm-Message-State: AOAM530KdN5wDKT2Oh7FtRT7/B5H6+iIDNkhoT/YfjbX+cH1ZTIVBURc dMpazEd8xgabKJ578ASG9wtPn95sSiUiI0lYhTiH7RzjgZMNGZX1878zYdcdq5EkDuCesCRrGqu nL+njjnhYPVsF58f9xFLeF2QFKJnsHcsbq4ApVzKb7IJGQwlIkvfx0/FPwmD5n8fr8zEBEkg= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:388b:9d0c:1bc4:40c]) (user=alankelly job=sendgmr) by 2002:a7b:c24a:: with SMTP id b10mr1164208wmj.191.1644398035435; Wed, 09 Feb 2022 01:13:55 -0800 (PST) Date: Wed, 9 Feb 2022 10:13:51 +0100 Message-Id: <20220209091351.3455295-1-alankelly@google.com> Mime-Version: 1.0 X-Mailer: git-send-email 2.35.0.263.gb82422642f-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 2/5] libswscale: Avx2 hscale can process inputs of any size. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: g3Ng3qJrcv6H The main loop processes blocks of 16 pixels. The tail processes blocks of size 4. --- libswscale/x86/scale_avx2.asm | 48 +++++++++++++++++++++++++++++++++-- 1 file changed, 46 insertions(+), 2 deletions(-) diff --git a/libswscale/x86/scale_avx2.asm b/libswscale/x86/scale_avx2.asm index 20acdbd633..dc42abb100 100644 --- a/libswscale/x86/scale_avx2.asm +++ b/libswscale/x86/scale_avx2.asm @@ -53,6 +53,9 @@ cglobal hscale8to15_%1, 7, 9, 16, pos0, dst, w, srcmem, filter, fltpos, fltsize, mova m14, [four] shr fltsized, 2 %endif + cmp wq, 16 + jl .tail_loop + mov countq, 0x10 .loop: movu m1, [fltposq] movu m2, [fltposq+32] @@ -97,11 +100,52 @@ cglobal hscale8to15_%1, 7, 9, 16, pos0, dst, w, srcmem, filter, fltpos, fltsize, vpsrad m6, 7 vpackssdw m5, m5, m6 vpermd m5, m15, m5 - vmovdqu [dstq + countq * 2], m5 + vmovdqu [dstq], m5 + add dstq, 0x20 add fltposq, 0x40 add countq, 0x10 cmp countq, wq - jl .loop + jle .loop + + sub countq, 0x10 + cmp countq, wq + jge .end + +.tail_loop: + movu xm1, [fltposq] +%ifidn %1, X4 + pxor xm9, xm9 + pxor xm10, xm10 + xor innerq, innerq +.tail_innerloop: +%endif + vpcmpeqd xm13, xm13 + vpgatherdd xm3,[srcmemq + xm1], xm13 + vpunpcklbw xm5, xm3, xm0 + vpunpckhbw xm6, xm3, xm0 + vpmaddwd xm5, xm5, [filterq] + vpmaddwd xm6, xm6, [filterq + 16] + add filterq, 0x20 +%ifidn %1, X4 + paddd xm9, xm5 + paddd xm10, xm6 + paddd xm1, xm14 + add innerq, 1 + cmp innerq, fltsizeq + jl .tail_innerloop + vphaddd xm5, xm9, xm10 +%else + vphaddd xm5, xm5, xm6 +%endif + vpsrad xm5, 7 + vpackssdw xm5, xm5, xm5 + vmovq [dstq], xm5 + add dstq, 0x8 + add fltposq, 0x10 + add countq, 0x4 + cmp countq, wq + jl .tail_loop +.end: REP_RET %endmacro From patchwork Wed Feb 9 09:14:17 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 34207 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6602:2c4e:0:0:0:0 with SMTP id x14csp1563381iov; Wed, 9 Feb 2022 01:14:32 -0800 (PST) X-Google-Smtp-Source: ABdhPJxOOUGx7bza4bUhkbp+8hLsK9Rht79UloOyjVE9b0dBjILqEZ+dl96jw/4EK/SufGdfEsU8 X-Received: by 2002:a17:907:9605:: with SMTP id gb5mr1155107ejc.490.1644398072140; Wed, 09 Feb 2022 01:14:32 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1644398072; cv=none; d=google.com; s=arc-20160816; b=Hb7gpfbXZVQxmCSkwvVhVsYFeYvjtVOqsNXMgpr8271vHZyPngKE7OtvoCZxDcmHQ1 lYMT+Wx+LZ4ggmPkyw80F5nwiQxPvbD8MZF3WKiQ5WGeqf4MfVfpidPm4B6M0AuSLQFi WGx7K5BcOziPj4bLxYI/5P596zpj4I7Hxna9CO1lbXaRVscfMRumgVhMSpLjpRYSM8Ss X2uJzX00xrt8G36n26vw6Sbh7UhTA+Ju4oLgdwsHjDSLOrX8sOLN+sDl+wLizzj2v+hW Y7agIqWw9Z48N8xEOuSM1xKwk+oOjNBgITuLPRdDdp1hggeOyDsh4iYCk7leJrd4dI4k cyPQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:mime-version:message-id:date :dkim-signature:delivered-to; bh=9Imf8d7ROLGHfLNvMVbSLJzZj/R77NO91T4F6uqoQOI=; b=mXF3vn4hLJrayIv5xjAHClqd7aqf+lpFnyG2nFHthUJIuOHClve1h4i0yiKfwgdWik kFmcEm9jiRdgcwppE4e2j+3GWbw8bT68oKN+wKn/ClY6bIkhm57lPv/N8we3bohGwaHX 7hTTx6ITQZ3m+LGpJ5Q7jxty6A5wtgI5PotDZ/a7lYQGv5MYEfxXpeQijLZuxbLS2f7H mQDeTgqwiTCZ0/s2lJ7PrgFyly+kGQl82MXLHsvrlt/62BO09DKb9xGmg+RiBvs+qjNj f2DBaI9kmegU8dV/oOh0Zm5DuT0iLJ0agd8l+OoEbUvP6tiVUKqjJLzzHh4UHUxrUbja DzoA== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=bAuoZROu; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id 2si10215249ejj.390.2022.02.09.01.14.31; Wed, 09 Feb 2022 01:14:32 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=bAuoZROu; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 21F6468AE72; Wed, 9 Feb 2022 11:14:29 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-wr1-f73.google.com (mail-wr1-f73.google.com [209.85.221.73]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id A27EE68AE72 for ; Wed, 9 Feb 2022 11:14:22 +0200 (EET) Received: by mail-wr1-f73.google.com with SMTP id e1-20020adfa741000000b001e2e74c3d4eso841601wrd.12 for ; Wed, 09 Feb 2022 01:14:22 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:message-id:mime-version:subject:from:to:cc; bh=Nmizo5e2bmRo9FbhWKdv9UYH3pHQCx2V2PBByeeGsxA=; b=bAuoZROuU3RNDfq8hT64AVOQkxwJuDi0YtcYwIZiD49EBsXCJ4D3CY1qgKxOJyz7NL +nmZji09efnPTEDviObF/ueGBaZkBdEvKxDZ8euuQKgxynx8mHFKuFeY/kiwMEPv9zNQ qbFr4UPafRPk3xEGEF2u/HjC+djKXY4VhXayPFc3TIPlPDddYO+jaK9hR8HJxFqmZ5In oZ9hN58eFiwbZeyRTU6Fz6EKXUVLwxCzInVHwELnt5UzS2zVv9vPeY4bwasIByvYa3SL 70gm7k5fQHh87/t+2T9LnN0/A45jOdQ5TOltrTK+I+xQWINYkQAYHQ+Z1LnzXlr60hUE zgQA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:message-id:mime-version:subject:from:to:cc; bh=Nmizo5e2bmRo9FbhWKdv9UYH3pHQCx2V2PBByeeGsxA=; b=MiRiVI1sQuW5xnTbJLN8uPzXXMkzR47WDQZ3ZT2KScBVgzp9m37dZj5+Zls/jl/HE5 ULRncdMNjikvT8MkPuMKgG3uoQfaZ40cpJ2jT2RnZSPmU6pAuEketT356D/TgvLQizZ3 VFuNN2AcxY7UfjyVbzqr+jBQ7K98yrpof2qeSKC8qepFmVkhutRa8Zubqmo4S9xW3+LY 3J5lQIw+PFj+k8kNFxHcLc+9qbA7nmr3meICkD9K2alnHe+aSqh61oh5AdlN6fnXHeSQ MNcmfsw43i03T0KvWip9fMQULbuh2RiOq+HIkRkLL1lbVMW/AAUhXKJ/h3OMgU+lmNU/ OX/w== X-Gm-Message-State: AOAM5336/czs7RnUcOtpElk4+cFmFblvyooAysS8HFfCbWxAmfVfNuwW brIZgfg5/48O9C4cQmPbSStHsFZ9QOF+B5bVFoYc7zwH+qHuHukf46gvEvQrxQWH6baC4ksTkdM xWERDXofSGZFn4jJ+6ymboFBTEfxy6nGGK53zqqXRfrwJhTgfhMkfAfkw68vuf/mot+AdmPg= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:388b:9d0c:1bc4:40c]) (user=alankelly job=sendgmr) by 2002:adf:de12:: with SMTP id b18mr1271964wrm.293.1644398061798; Wed, 09 Feb 2022 01:14:21 -0800 (PST) Date: Wed, 9 Feb 2022 10:14:17 +0100 Message-Id: <20220209091417.3456063-1-alankelly@google.com> Mime-Version: 1.0 X-Mailer: git-send-email 2.35.0.263.gb82422642f-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 3/5] libswscale: Enable hscale_avx2 for all input sizes. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: wY3EklV4CaDU ff_shuffle_filter_coefficients shuffles the tail as required. --- libswscale/utils.c | 19 ++++++++++++++++--- libswscale/x86/swscale.c | 6 ++---- 2 files changed, 18 insertions(+), 7 deletions(-) diff --git a/libswscale/utils.c b/libswscale/utils.c index 1d919e863a..31c365fcee 100644 --- a/libswscale/utils.c +++ b/libswscale/utils.c @@ -285,8 +285,7 @@ void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, #if ARCH_X86_64 int i, j, k; int cpu_flags = av_get_cpu_flags(); - // avx2 hscale filter processes 16 pixel blocks. - if (!filter || dstW % 16 != 0) + if (!filter) return; if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { @@ -298,9 +297,11 @@ void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } // Do not swap filterPos for pixels which won't be processed by // the main loop. - for (i = 0; i + 8 <= dstW; i += 8) { + for (i = 0; i + 16 <= dstW; i += 16) { FFSWAP(int, filterPos[i + 2], filterPos[i + 4]); FFSWAP(int, filterPos[i + 3], filterPos[i + 5]); + FFSWAP(int, filterPos[i + 10], filterPos[i + 12]); + FFSWAP(int, filterPos[i + 11], filterPos[i + 13]); } if (filterSize > 4) { // 16 pixels are processed at a time. @@ -314,6 +315,18 @@ void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } } } + // 4 pixels are processed at a time in the tail. + for (; i < dstW; i += 4) { + // 4 filter coeffs are processed at a time. + int rem = dstW - i >= 4 ? 4 : dstW - i; + for (k = 0; k + 4 <= filterSize; k += 4) { + for (j = 0; j < rem; ++j) { + int from = (i + j) * filterSize + k; + int to = i * filterSize + j * 4 + k * 4; + memcpy(&filter[to], &filterCopy[from], 4 * sizeof(int16_t)); + } + } + } } if (filterCopy) av_free(filterCopy); diff --git a/libswscale/x86/swscale.c b/libswscale/x86/swscale.c index 73869355b8..76f5a70fc5 100644 --- a/libswscale/x86/swscale.c +++ b/libswscale/x86/swscale.c @@ -691,10 +691,8 @@ switch(c->dstBpc){ \ if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { - if (c->chrDstW % 16 == 0) - ASSIGN_AVX2_SCALE_FUNC(c->hcScale, c->hChrFilterSize); - if (c->dstW % 16 == 0) - ASSIGN_AVX2_SCALE_FUNC(c->hyScale, c->hLumFilterSize); + ASSIGN_AVX2_SCALE_FUNC(c->hcScale, c->hChrFilterSize); + ASSIGN_AVX2_SCALE_FUNC(c->hyScale, c->hLumFilterSize); } } From patchwork Wed Feb 9 09:14:42 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 34208 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6602:2c4e:0:0:0:0 with SMTP id x14csp1563679iov; Wed, 9 Feb 2022 01:14:55 -0800 (PST) X-Google-Smtp-Source: ABdhPJyogANbtPb4S2lb33jv9MnjOPuLR8h6SWsiQOcyF+jQl4mAYBobEcaYTQrT2jZ/8+4S7Pvr X-Received: by 2002:a17:906:7314:: with SMTP id di20mr510206ejc.259.1644398095246; Wed, 09 Feb 2022 01:14:55 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1644398095; cv=none; d=google.com; s=arc-20160816; b=XxBSWbBfYqLN8yNe6UVMl0GUbh5gj+Vs2vUmXRVYhqFJqncB0voSBb1vR/7Fg9l12L ucjeDTcxywV69Tp79CAFcRAz01vOVxPS9DyFFqGZIBAgQfktsLvZPso/3xSCov/EX2SW qokFyyCVAqbUTXzkvJwLQ3KrIeM5cTZiI7VHe29JOKKv0ZmdeCrpLj8FJpkr6n0hIrDb Io/w9ul3MwIocV4O5U1sKf/Hb3yI+iuL/a/bXP2lzOZ58n24C+Mro7xZrh0ezmoxwUBr CZ8QRNYNfuE/GYVQS96YPI8XxyZfLHu/gC7KQh6qI2QTRwh7Jar1GrQ2eJVmI2ptV1dX ujLA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:mime-version:message-id:date :dkim-signature:delivered-to; bh=l7zuD1XQNL8oCnaPYo+f/PaucAunEd9UWb7g0HSg83M=; b=epjm8I7/3heaqMnc6BhSKVnC8KaVq3YgWVPMdyAn58goZgiE7n6v5LfrhUd1JpIVws +hBmUV3nvzkpLuroTbvli5WiWDhnAJ6kRvtif1S+jPx6Y+KWgG0F3VsOjnBLbhTcZpeQ PWsH19Vps7VV044hv7BsQmnMsV307jsQAO4Z9HjclR0kv30cEy+Dn483PaS6Il9wssAW snb4JIA38Y8Swt1yX0ktV9mMSWHW3KOuLIUenv4yUUj0qEXXrP6nBxd4ZvSzLfn7G905 TGBx/CzqRcFd4szYt+xc1pGgdDGtIRp/pfBK5qfESguxx8KNclF+dfJGwZogiTWdTEbr kyag== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=oENRHJQU; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id z13si12768031edd.111.2022.02.09.01.14.54; Wed, 09 Feb 2022 01:14:55 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=oENRHJQU; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 2F09568B1F8; Wed, 9 Feb 2022 11:14:52 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-wr1-f74.google.com (mail-wr1-f74.google.com [209.85.221.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id D503068B1CA for ; Wed, 9 Feb 2022 11:14:45 +0200 (EET) Received: by mail-wr1-f74.google.com with SMTP id w26-20020adf8bda000000b001e33dbc525cso829250wra.18 for ; Wed, 09 Feb 2022 01:14:45 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:message-id:mime-version:subject:from:to:cc; bh=xmp3gK13oFSRUDn8rWg1daSwCfIJwY0N+S2UW+Nr4uA=; b=oENRHJQU634by+9yLln0ujsxHTR2ecg6iw0eVZGvATJDAJJPTQUgUKd7Pd112S/ofN /qx3aXb0pwYXnCkGN1xleRUv2k8HA+9hMpkr+5/qp73NeT/qwGBbm/KSdMqLMX9vSljE o0lo4rPsGQJEwbkuydGCwT5GqJOoLlIq4CflJTsgZHrDN5jppAM4eDdL5PoUxFZ1bCoP zqbxXJ+t7xOcViPPSdYUUXP28Uez0N7LFP5zcdq3xZDjIA8KvwmnWrWhY/v7oT2ODfwb Fz1oR70PYZvGfZu41Ca2+g7GGCrMGae+APBH8LrtO8LMPcNCyB94GKLamrJgpPrlfxNq /o1w== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:message-id:mime-version:subject:from:to:cc; bh=xmp3gK13oFSRUDn8rWg1daSwCfIJwY0N+S2UW+Nr4uA=; b=OLzfedJNlZDC8DoxaILTVD8enJmyoJ7glE8RL8hrcHUwMGCFEQrFXAGNC/f7+Qix6B i8vQBSBgUMwNdktRJHENaXjxccWbxzT5Jlru+mGErCfw/clun2lJPFpivqljIiIVG51G Gqtkxuf+x9/1vC9g00OHaT8Bt49074wubmlha+x4KmBJOXpRQwDdxm1UBDFYyHsyCyPG nYwgIzCo8633rZnuZwQ5SNyZf+IJnBhiHr4EU/8bKY4YIW3VljphfoGvo3Ni1DcAcFrk FJMX3S7Eg5Dy9o0Pxy39GsRxFIYq9zjTDmShVOVHPzioMcp+Eok6fAzYv0LlW858OcaF Q09w== X-Gm-Message-State: AOAM5330wS2LPScO7DzdQ0QFyHqKzTcfxpBpnDqu0w6fO42e7u6Egb9R 4T8tduw38+H5p7cEGVtcTBRz0JDpcMMnJaImh/zdkP/09M6jIZ/9Y8S4H7ZA49reowvwux5ZNiU oQ3sCv114ctP6Fbcow9TSCsv+WuQTY+SiywAVvvazdWi3h7Gb1D+vuXIiMJP2uXbaJtbtTHE= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:388b:9d0c:1bc4:40c]) (user=alankelly job=sendgmr) by 2002:a1c:ed12:: with SMTP id l18mr1682515wmh.93.1644398085277; Wed, 09 Feb 2022 01:14:45 -0800 (PST) Date: Wed, 9 Feb 2022 10:14:42 +0100 Message-Id: <20220209091442.3456781-1-alankelly@google.com> Mime-Version: 1.0 X-Mailer: git-send-email 2.35.0.263.gb82422642f-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 4/5] libswscale: Propagate error codes from ff_shuffle_filter_coefficients X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: c9Nh7u6LlsnZ --- libswscale/swscale_internal.h | 2 +- libswscale/utils.c | 14 ++++++++------ 2 files changed, 9 insertions(+), 7 deletions(-) diff --git a/libswscale/swscale_internal.h b/libswscale/swscale_internal.h index 3a78d95ba6..26d28d42e6 100644 --- a/libswscale/swscale_internal.h +++ b/libswscale/swscale_internal.h @@ -1144,5 +1144,5 @@ void ff_sws_slice_worker(void *priv, int jobnr, int threadnr, #define MAX_LINES_AHEAD 4 //shuffle filter and filterPos for hyScale and hcScale filters in avx2 -void ff_shuffle_filter_coefficients(SwsContext *c, int* filterPos, int filterSize, int16_t *filter, int dstW); +int ff_shuffle_filter_coefficients(SwsContext *c, int* filterPos, int filterSize, int16_t *filter, int dstW); #endif /* SWSCALE_SWSCALE_INTERNAL_H */ diff --git a/libswscale/utils.c b/libswscale/utils.c index 31c365fcee..1f8705a417 100644 --- a/libswscale/utils.c +++ b/libswscale/utils.c @@ -278,7 +278,7 @@ static const FormatEntry format_entries[] = { [AV_PIX_FMT_P416LE] = { 1, 1 }, }; -void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, +int ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, int filterSize, int16_t *filter, int dstW) { @@ -286,13 +286,13 @@ void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, int i, j, k; int cpu_flags = av_get_cpu_flags(); if (!filter) - return; + return 0; if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { int16_t *filterCopy = NULL; if (filterSize > 4) { if (!FF_ALLOC_TYPED_ARRAY(filterCopy, dstW * filterSize)) - return; + return AVERROR(ENOMEM); memcpy(filterCopy, filter, dstW * filterSize * sizeof(int16_t)); } // Do not swap filterPos for pixels which won't be processed by @@ -333,7 +333,7 @@ void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } } #endif - return; + return 0; } int sws_isSupportedInput(enum AVPixelFormat pix_fmt) @@ -1859,7 +1859,8 @@ av_cold int sws_init_context(SwsContext *c, SwsFilter *srcFilter, get_local_pos(c, 0, 0, 0), get_local_pos(c, 0, 0, 0))) < 0) goto fail; - ff_shuffle_filter_coefficients(c, c->hLumFilterPos, c->hLumFilterSize, c->hLumFilter, dstW); + if ((ff_shuffle_filter_coefficients(c, c->hLumFilterPos, c->hLumFilterSize, c->hLumFilter, dstW)) < 0) + goto nomem; if ((ret = initFilter(&c->hChrFilter, &c->hChrFilterPos, &c->hChrFilterSize, c->chrXInc, c->chrSrcW, c->chrDstW, filterAlign, 1 << 14, @@ -1869,7 +1870,8 @@ av_cold int sws_init_context(SwsContext *c, SwsFilter *srcFilter, get_local_pos(c, c->chrSrcHSubSample, c->src_h_chr_pos, 0), get_local_pos(c, c->chrDstHSubSample, c->dst_h_chr_pos, 0))) < 0) goto fail; - ff_shuffle_filter_coefficients(c, c->hChrFilterPos, c->hChrFilterSize, c->hChrFilter, c->chrDstW); + if ((ff_shuffle_filter_coefficients(c, c->hChrFilterPos, c->hChrFilterSize, c->hChrFilter, c->chrDstW)) < 0) + goto nomem; } } // initialize horizontal stuff From patchwork Wed Feb 9 09:15:09 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 34209 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6602:2c4e:0:0:0:0 with SMTP id x14csp1563886iov; Wed, 9 Feb 2022 01:15:20 -0800 (PST) X-Google-Smtp-Source: ABdhPJwQOZie27yB2FMT/+tRs8VlrqkVa0TVBmiCh++bApIHcR+SpKq+TV2N6xqJANtW8AnBw3pc X-Received: by 2002:a17:907:1c04:: with SMTP id nc4mr1064586ejc.370.1644398119880; Wed, 09 Feb 2022 01:15:19 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1644398119; cv=none; d=google.com; s=arc-20160816; b=mVqn4yKfW4DBhB6NxzwumtgnEff2MRyYO8oHFBHMEhD8Olhn/JpfiuYkWxylWmRd1R ibna+uWFxKq3lRkKBK40jFX7Fxb/cRkiZY6e5H3Kn61xgIsNb322JnXrpe9ixYlhaLWG Tp3R3PrkZuKxXVl3wmi4cant/qzReGTO8BWS7AiQrj2dgtWwaUVJp0nHfrCj7d5NOQNe chUSbukmIdJjoAabIvT4pYXuGHzB0g3pIQ+C95qmn6ZEQS649MBVwUfHLBYzopEWTB6k x8Pt2NcewFEPAVZYG+q/0ZHjTcid++cs30LJdO3ZXSXB5jE/CH0+a8GKJcDxDOLEmcsH mXbw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:mime-version:message-id:date :dkim-signature:delivered-to; bh=HTKLMCeigxVQbx0iUB99jwqP1eLVEPRt6lP1+8zv1go=; b=1J+zEFk5WF5kXXwPdfN9N/5Lu4Ki7gzqqT53MyL/lOIPCoV/9IgUa3eXd1EfY4Wn8+ JQEZVJddQ7fTpt962Uv4XO2jpuuzlOz3e/KQVnrytKLe6vIFSzRmA8hj02lOFbxahzcf RuWx61WWV/xj/MQ5Qx+x+x0TIdx9jZ1vfROG7QhdR8BemmYDcr/jTKENbq8yu78yc4xD MEmpqKSoJRJCpWaOTeyWa/tNt9L/Qs8e3q4GxLy7slbbZilWFzykZHfSbmPw61VYGgSH 1yxd4Wdo+A4NgQtTOULAIHSKQ4z4fdOraam0VfgOfhL4Y+P2zswO7cDR0orraqzeyFCH TRDA== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=BF5f1Bdo; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id jg3si4530001ejc.83.2022.02.09.01.15.19; Wed, 09 Feb 2022 01:15:19 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=BF5f1Bdo; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 3FF5768B1F9; Wed, 9 Feb 2022 11:15:17 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-wr1-f74.google.com (mail-wr1-f74.google.com [209.85.221.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id D52D968B1E0 for ; Wed, 9 Feb 2022 11:15:13 +0200 (EET) Received: by mail-wr1-f74.google.com with SMTP id n18-20020adfc612000000b001e3310ca453so855126wrg.2 for ; Wed, 09 Feb 2022 01:15:13 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:message-id:mime-version:subject:from:to:cc; bh=N7arwcUCOtZk0lD1h9i96kmmKYg3QQFgxH39iNxCsRY=; b=BF5f1BdoIr062ngvCMVbfufSbc4+TSM9kEoLE7R5YV+VBb94xNSd4xCy3/u5WYoMb7 VZ6tBHbMLQ0VPstmdrcbV6uI2zIm2+NZ3uDdtN6JQPgJMXIdM42pbsFI00tJuskxn82b 84PayDa6Y/UaSs80lPAXezNo/kDQleBhk21DaaDG4O3kAdsG0pzAQbbT7wyJMi0KW4RM nSe/GubMMSkvWQI+cZ69oHW4aX7JZ1t1a0rktv9VODT3feghbpwYQBFaqJTv3SHBFKkf T9AvLxg3ajgnacthih9lf6JDcXtsUAjDrU/3ZkFGgABcD4kNT1clSWZ45MqQo++hOIkR 38LA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:message-id:mime-version:subject:from:to:cc; bh=N7arwcUCOtZk0lD1h9i96kmmKYg3QQFgxH39iNxCsRY=; b=NuQ/cxL1ivd1FHXEwqJeE9KuX21ORC1v5AB80YofZxWdHYWxvD3zsurqMsBEF4UJlr L3M2FhOPj44mdIZuL5TXJ3hw/Y+k+vzW8D1GE4CziQYwZWG8ZlAWzEI3daMl70ADyfkJ Z/6ThVdpqTXh90zdVqudUL+6gr73aLYUU7tS8JTvDSl7PcjepGzq7+2AwZUTECKhbi/q c7fUqIvn2vnQMliZDKdxNbv9qJIt9xVMOHLTETtmQgyAo4cH/2gE6XHpcuvjGb3Ts6CX 9fFowN06ON4Buy8oeWVGMo4D92BsaWripAqH9EdZsQ179rbkYxmZwBRJz3SJEjiAPBMD uMuA== X-Gm-Message-State: AOAM531HnLG0F4aHuBup9zV2yTXJ15fYtU3bX+YnOogbKoh5McVXCZ/c zJlgK8B5/vesYpdt44CezmcFCKbQDK7zohP5YpL3KSbbwbzbG0Tq6SLm/+vTS/Iq4wRuEzuGFua u0uBkpUB6fIRY6YQSDzwKau9fj8Ym2QK4mwcdAP0qty/FK4nUs/gQ4iswIBZOgYew2AAr/VQ= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:388b:9d0c:1bc4:40c]) (user=alankelly job=sendgmr) by 2002:a05:600c:290c:: with SMTP id i12mr1172705wmd.95.1644398112972; Wed, 09 Feb 2022 01:15:12 -0800 (PST) Date: Wed, 9 Feb 2022 10:15:09 +0100 Message-Id: <20220209091509.3457523-1-alankelly@google.com> Mime-Version: 1.0 X-Mailer: git-send-email 2.35.0.263.gb82422642f-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 5/5] checkasm/sw_scale: hscale does not requires cpuflag test. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: UpC2Uq6CB9Ds This is done in ff_shuffle_filter_coefficients. --- tests/checkasm/sw_scale.c | 6 ++---- 1 file changed, 2 insertions(+), 4 deletions(-) diff --git a/tests/checkasm/sw_scale.c b/tests/checkasm/sw_scale.c index 3c0a083b42..e7f916d3a8 100644 --- a/tests/checkasm/sw_scale.c +++ b/tests/checkasm/sw_scale.c @@ -168,8 +168,6 @@ static void check_hscale(void) const uint8_t *src, const int16_t *filter, const int32_t *filterPos, int filterSize); - int cpu_flags = av_get_cpu_flags(); - ctx = sws_alloc_context(); if (sws_init_context(ctx, NULL, NULL) < 0) fail(); @@ -215,10 +213,10 @@ static void check_hscale(void) filter[SRC_PIXELS * width + i] = rnd(); } + ff_sws_init_scale(ctx); memcpy(filterAvx2, filter, sizeof(uint16_t) * (SRC_PIXELS * MAX_FILTER_WIDTH + MAX_FILTER_WIDTH)); - if ((cpu_flags & AV_CPU_FLAG_AVX2) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) - ff_shuffle_filter_coefficients(ctx, filterPosAvx, width, filterAvx2, SRC_PIXELS); + ff_shuffle_filter_coefficients(ctx, filterPosAvx, width, filterAvx2, SRC_PIXELS); if (check_func(ctx->hcScale, "hscale_%d_to_%d_width%d", ctx->srcBpc, ctx->dstBpc + 1, width)) { memset(dst0, 0, SRC_PIXELS * sizeof(dst0[0]));