From patchwork Wed Feb 9 09:14:17 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 34207 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6602:2c4e:0:0:0:0 with SMTP id x14csp1563381iov; Wed, 9 Feb 2022 01:14:32 -0800 (PST) X-Google-Smtp-Source: ABdhPJxOOUGx7bza4bUhkbp+8hLsK9Rht79UloOyjVE9b0dBjILqEZ+dl96jw/4EK/SufGdfEsU8 X-Received: by 2002:a17:907:9605:: with SMTP id gb5mr1155107ejc.490.1644398072140; Wed, 09 Feb 2022 01:14:32 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1644398072; cv=none; d=google.com; s=arc-20160816; b=Hb7gpfbXZVQxmCSkwvVhVsYFeYvjtVOqsNXMgpr8271vHZyPngKE7OtvoCZxDcmHQ1 lYMT+Wx+LZ4ggmPkyw80F5nwiQxPvbD8MZF3WKiQ5WGeqf4MfVfpidPm4B6M0AuSLQFi WGx7K5BcOziPj4bLxYI/5P596zpj4I7Hxna9CO1lbXaRVscfMRumgVhMSpLjpRYSM8Ss X2uJzX00xrt8G36n26vw6Sbh7UhTA+Ju4oLgdwsHjDSLOrX8sOLN+sDl+wLizzj2v+hW Y7agIqWw9Z48N8xEOuSM1xKwk+oOjNBgITuLPRdDdp1hggeOyDsh4iYCk7leJrd4dI4k cyPQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:mime-version:message-id:date :dkim-signature:delivered-to; bh=9Imf8d7ROLGHfLNvMVbSLJzZj/R77NO91T4F6uqoQOI=; b=mXF3vn4hLJrayIv5xjAHClqd7aqf+lpFnyG2nFHthUJIuOHClve1h4i0yiKfwgdWik kFmcEm9jiRdgcwppE4e2j+3GWbw8bT68oKN+wKn/ClY6bIkhm57lPv/N8we3bohGwaHX 7hTTx6ITQZ3m+LGpJ5Q7jxty6A5wtgI5PotDZ/a7lYQGv5MYEfxXpeQijLZuxbLS2f7H mQDeTgqwiTCZ0/s2lJ7PrgFyly+kGQl82MXLHsvrlt/62BO09DKb9xGmg+RiBvs+qjNj f2DBaI9kmegU8dV/oOh0Zm5DuT0iLJ0agd8l+OoEbUvP6tiVUKqjJLzzHh4UHUxrUbja DzoA== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=bAuoZROu; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id 2si10215249ejj.390.2022.02.09.01.14.31; Wed, 09 Feb 2022 01:14:32 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=bAuoZROu; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 21F6468AE72; Wed, 9 Feb 2022 11:14:29 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-wr1-f73.google.com (mail-wr1-f73.google.com [209.85.221.73]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id A27EE68AE72 for ; Wed, 9 Feb 2022 11:14:22 +0200 (EET) Received: by mail-wr1-f73.google.com with SMTP id e1-20020adfa741000000b001e2e74c3d4eso841601wrd.12 for ; Wed, 09 Feb 2022 01:14:22 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:message-id:mime-version:subject:from:to:cc; bh=Nmizo5e2bmRo9FbhWKdv9UYH3pHQCx2V2PBByeeGsxA=; b=bAuoZROuU3RNDfq8hT64AVOQkxwJuDi0YtcYwIZiD49EBsXCJ4D3CY1qgKxOJyz7NL +nmZji09efnPTEDviObF/ueGBaZkBdEvKxDZ8euuQKgxynx8mHFKuFeY/kiwMEPv9zNQ qbFr4UPafRPk3xEGEF2u/HjC+djKXY4VhXayPFc3TIPlPDddYO+jaK9hR8HJxFqmZ5In oZ9hN58eFiwbZeyRTU6Fz6EKXUVLwxCzInVHwELnt5UzS2zVv9vPeY4bwasIByvYa3SL 70gm7k5fQHh87/t+2T9LnN0/A45jOdQ5TOltrTK+I+xQWINYkQAYHQ+Z1LnzXlr60hUE zgQA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:message-id:mime-version:subject:from:to:cc; bh=Nmizo5e2bmRo9FbhWKdv9UYH3pHQCx2V2PBByeeGsxA=; b=MiRiVI1sQuW5xnTbJLN8uPzXXMkzR47WDQZ3ZT2KScBVgzp9m37dZj5+Zls/jl/HE5 ULRncdMNjikvT8MkPuMKgG3uoQfaZ40cpJ2jT2RnZSPmU6pAuEketT356D/TgvLQizZ3 VFuNN2AcxY7UfjyVbzqr+jBQ7K98yrpof2qeSKC8qepFmVkhutRa8Zubqmo4S9xW3+LY 3J5lQIw+PFj+k8kNFxHcLc+9qbA7nmr3meICkD9K2alnHe+aSqh61oh5AdlN6fnXHeSQ MNcmfsw43i03T0KvWip9fMQULbuh2RiOq+HIkRkLL1lbVMW/AAUhXKJ/h3OMgU+lmNU/ OX/w== X-Gm-Message-State: AOAM5336/czs7RnUcOtpElk4+cFmFblvyooAysS8HFfCbWxAmfVfNuwW brIZgfg5/48O9C4cQmPbSStHsFZ9QOF+B5bVFoYc7zwH+qHuHukf46gvEvQrxQWH6baC4ksTkdM xWERDXofSGZFn4jJ+6ymboFBTEfxy6nGGK53zqqXRfrwJhTgfhMkfAfkw68vuf/mot+AdmPg= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:388b:9d0c:1bc4:40c]) (user=alankelly job=sendgmr) by 2002:adf:de12:: with SMTP id b18mr1271964wrm.293.1644398061798; Wed, 09 Feb 2022 01:14:21 -0800 (PST) Date: Wed, 9 Feb 2022 10:14:17 +0100 Message-Id: <20220209091417.3456063-1-alankelly@google.com> Mime-Version: 1.0 X-Mailer: git-send-email 2.35.0.263.gb82422642f-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 3/5] libswscale: Enable hscale_avx2 for all input sizes. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: wY3EklV4CaDU ff_shuffle_filter_coefficients shuffles the tail as required. --- libswscale/utils.c | 19 ++++++++++++++++--- libswscale/x86/swscale.c | 6 ++---- 2 files changed, 18 insertions(+), 7 deletions(-) diff --git a/libswscale/utils.c b/libswscale/utils.c index 1d919e863a..31c365fcee 100644 --- a/libswscale/utils.c +++ b/libswscale/utils.c @@ -285,8 +285,7 @@ void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, #if ARCH_X86_64 int i, j, k; int cpu_flags = av_get_cpu_flags(); - // avx2 hscale filter processes 16 pixel blocks. - if (!filter || dstW % 16 != 0) + if (!filter) return; if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { @@ -298,9 +297,11 @@ void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } // Do not swap filterPos for pixels which won't be processed by // the main loop. - for (i = 0; i + 8 <= dstW; i += 8) { + for (i = 0; i + 16 <= dstW; i += 16) { FFSWAP(int, filterPos[i + 2], filterPos[i + 4]); FFSWAP(int, filterPos[i + 3], filterPos[i + 5]); + FFSWAP(int, filterPos[i + 10], filterPos[i + 12]); + FFSWAP(int, filterPos[i + 11], filterPos[i + 13]); } if (filterSize > 4) { // 16 pixels are processed at a time. @@ -314,6 +315,18 @@ void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, } } } + // 4 pixels are processed at a time in the tail. + for (; i < dstW; i += 4) { + // 4 filter coeffs are processed at a time. + int rem = dstW - i >= 4 ? 4 : dstW - i; + for (k = 0; k + 4 <= filterSize; k += 4) { + for (j = 0; j < rem; ++j) { + int from = (i + j) * filterSize + k; + int to = i * filterSize + j * 4 + k * 4; + memcpy(&filter[to], &filterCopy[from], 4 * sizeof(int16_t)); + } + } + } } if (filterCopy) av_free(filterCopy); diff --git a/libswscale/x86/swscale.c b/libswscale/x86/swscale.c index 73869355b8..76f5a70fc5 100644 --- a/libswscale/x86/swscale.c +++ b/libswscale/x86/swscale.c @@ -691,10 +691,8 @@ switch(c->dstBpc){ \ if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { - if (c->chrDstW % 16 == 0) - ASSIGN_AVX2_SCALE_FUNC(c->hcScale, c->hChrFilterSize); - if (c->dstW % 16 == 0) - ASSIGN_AVX2_SCALE_FUNC(c->hyScale, c->hLumFilterSize); + ASSIGN_AVX2_SCALE_FUNC(c->hcScale, c->hChrFilterSize); + ASSIGN_AVX2_SCALE_FUNC(c->hyScale, c->hLumFilterSize); } }