From patchwork Wed Jun 5 20:28:52 2024 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: James Almer X-Patchwork-Id: 49603 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a59:9185:0:b0:460:55fa:d5ed with SMTP id s5csp11688vqg; Wed, 5 Jun 2024 13:29:21 -0700 (PDT) X-Forwarded-Encrypted: i=2; AJvYcCUZfkkp/Re/1n5F/NqnRBBuSyF6iZUJovNv35DvFS+0aP9wCgQkFXmp5TsIG8K0Zbh0oSaKqR3mq3DOOJaJlM7nC/4Yor7YW6MrTw== X-Google-Smtp-Source: AGHT+IHn4WL8xs975hYPhvKAgtishpQFIcSMgPjUZhAoW7TB4lmPNQqzYpGE1XI1L6SZVAhIK41S X-Received: by 2002:a2e:92d6:0:b0:2de:d4ef:af19 with SMTP id 38308e7fff4ca-2eac79ba70dmr22528561fa.10.1717619360846; Wed, 05 Jun 2024 13:29:20 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1717619360; cv=none; d=google.com; s=arc-20160816; b=d9xds1eo4o9kldcmBLg6RVMPeYc2tp/61EcWOk3LrMDhKKJQ+EP10gMKqonsVboHn+ uFuIWDs381XwxpFtleWdTMiZ1di5cY2FZE8JHq936uJ8FFdTcdXY+PWSpu3Ny10uVt0A yZjk/JZCi0cV/wSl/bV8TGmmf/6W0VbFuDV5ZMCc9Ri0zjlZcxxnWm0IylXimBPY9OIo aDV0Qpa+7ED8Vfk6u9j+pwXosLVt2od6hXSrLVUAE2XjDjqJhzcmxHhDaUfEFwoBuG/A FlHQBoeDZYSvme/PZqNx3ga5Kc+AOUeswHaGi4ZkMovTnBi4hb4xjJrfO3FqL8j9XLbg t+jg== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:reply-to:list-subscribe :list-help:list-post:list-archive:list-unsubscribe:list-id :precedence:subject:mime-version:message-id:date:to:from :dkim-signature:delivered-to; bh=TPio1cEbhYzWM14NVsZ1MSACFPpLF2d0jWNLxMbbTfM=; fh=YOA8vD9MJZuwZ71F/05pj6KdCjf6jQRmzLS+CATXUQk=; b=Ef4fEEHczCofGgIFNVrTaLSFaRz3m7WsTnV5zVzB9uHb3cpCoclHar0Xgls0NmWUWB 5iNx5QN9T6MD0lbb4JVYSjMc2XL8iCvfaNpuA6NVUM4SpkhqxRXHkTRVmZRdUnnnHTVb nOgEy9t/lbDg9fa5msmTxJjGrIzoJdR+GCXbddrEs/PXy7pXFKwPfH/11uPqPdhcclJG fOd7QTVR/pXQ1fHV/ryrIaYx1wmx0VKJsDK2emIwxaCa1k4sS87bTJc9gVqXYl+IIQmv cVLOuDa+692FHeFsXvdE8rppojWJJb1RmodHo4V6OlzAn4S8KF7HwSOqSmy+arQVBa2C kWKA==; dara=google.com ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@gmail.com header.s=20230601 header.b=WFpeWqr8; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=QUARANTINE dis=NONE) header.from=gmail.com Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id 4fb4d7f45d1cf-57a31b90506si6445840a12.81.2024.06.05.13.29.20; Wed, 05 Jun 2024 13:29:20 -0700 (PDT) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@gmail.com header.s=20230601 header.b=WFpeWqr8; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=QUARANTINE dis=NONE) header.from=gmail.com Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 2CA3E68D6B2; Wed, 5 Jun 2024 23:29:18 +0300 (EEST) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-pf1-f171.google.com (mail-pf1-f171.google.com [209.85.210.171]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 531B268D44A for ; Wed, 5 Jun 2024 23:29:12 +0300 (EEST) Received: by mail-pf1-f171.google.com with SMTP id d2e1a72fcca58-6f4603237e0so132517b3a.0 for ; Wed, 05 Jun 2024 13:29:12 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1717619350; x=1718224150; darn=ffmpeg.org; h=content-transfer-encoding:mime-version:message-id:date:subject:to :from:from:to:cc:subject:date:message-id:reply-to; bh=tGXUtDF3Qhy7W7MfaKgPefGOtv74hS5wdJz98O26QlU=; b=WFpeWqr8+qfMgUzzx8qMDubwuC6M/ZYQM4TogFmYo85xhWlvAgZ2J/SI4JSfF9PZDf pMMGAg8EZFTQ/75GUlAQv6MoQLH7mL6p/UuTf3W8LowRL6s1uTFerNzYBEeSPr9EtAMs YuatZ+Z2jD7V4aOIrq+7KpYaGF3mFYdpGy37ecfDKYULYCMjelMJvQ6E7wYFMUE0vKSf KInpVwzKUKFl4Lawr2Wq/kAZoKMOg/IRnxY91Cbpo8N/CbANIUd6FCcs7aeAfr2V1ja6 tsS/irpyX3ScSYWbNt+cOMizyLlh+CyCbv2bocGv4/SOBULoiri8TuNFpCF0/q0jKlTV X3ig== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1717619350; x=1718224150; h=content-transfer-encoding:mime-version:message-id:date:subject:to :from:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=tGXUtDF3Qhy7W7MfaKgPefGOtv74hS5wdJz98O26QlU=; b=wAOuwl+komsyyEkQCBLqZsAkzz2noCccfrmRLrMnz64ecBjoIIbQnohSAJWKRPcNqI XxS+nCWYs470R0Xq4ycEWKgtMPi8HHdZHFXMAioHfqJpu2I+l9soJZNEznk59PKsPzDi J8IgzL8wiePWs5XBzwEq8HiJ8djku9ZZw0r1Z2WTl4pLn7erh5kA2vsTtdib3vEBowZX y95eM5RK0Db9E58DrD3nceLMwS+qbqWkgscy5wwkBpNq+vwK6FAN7+cnZ4+NOdEQjnyR 0D/S1q8UUIGHSU83DJydz5uW8OfbN63XA7RkAfuQl2bEoIr1IKxcTuoKkLnAOYUCcOne 2DwQ== X-Gm-Message-State: AOJu0YzYjSbucsEqmOz8xEbT9VLlwLEcywSCZuW8MY/piqIEHXxTfEgT le9VugsyC79tzpX0syEgPhxu2PHa7NjWT6Au7gZGBng/frHYgh2sBJMt3w== X-Received: by 2002:a05:6a00:2d83:b0:6f3:8468:8bb with SMTP id d2e1a72fcca58-703f88a8a51mr905163b3a.17.1717619349895; Wed, 05 Jun 2024 13:29:09 -0700 (PDT) Received: from localhost.localdomain ([190.194.167.233]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-7024967ae05sm8536492b3a.157.2024.06.05.13.29.08 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 05 Jun 2024 13:29:09 -0700 (PDT) From: James Almer To: ffmpeg-devel@ffmpeg.org Date: Wed, 5 Jun 2024 17:28:52 -0300 Message-ID: <20240605202853.3135-1-jamrial@gmail.com> X-Mailer: git-send-email 2.45.1 MIME-Version: 1.0 Subject: [FFmpeg-devel] [PATCH 1/2] swscale/x86/input: add AVX2 optimized RGB32 to YUV functions X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: wCwiLLkWzycF abgr_to_uv_8_c: 43.3 abgr_to_uv_8_sse2: 14.3 abgr_to_uv_8_avx: 15.3 abgr_to_uv_8_avx2: 18.8 abgr_to_uv_128_c: 650.3 abgr_to_uv_128_sse2: 110.8 abgr_to_uv_128_avx: 112.3 abgr_to_uv_128_avx2: 64.8 abgr_to_uv_1080_c: 5456.3 abgr_to_uv_1080_sse2: 888.8 abgr_to_uv_1080_avx: 900.8 abgr_to_uv_1080_avx2: 518.3 abgr_to_uv_1920_c: 9692.3 abgr_to_uv_1920_sse2: 1593.8 abgr_to_uv_1920_avx: 1613.3 abgr_to_uv_1920_avx2: 864.8 abgr_to_y_8_c: 23.3 abgr_to_y_8_sse2: 12.8 abgr_to_y_8_avx: 13.3 abgr_to_y_8_avx2: 17.3 abgr_to_y_128_c: 308.3 abgr_to_y_128_sse2: 67.3 abgr_to_y_128_avx: 66.8 abgr_to_y_128_avx2: 44.8 abgr_to_y_1080_c: 2371.3 abgr_to_y_1080_sse2: 512.8 abgr_to_y_1080_avx: 505.8 abgr_to_y_1080_avx2: 314.3 abgr_to_y_1920_c: 4177.3 abgr_to_y_1920_sse2: 915.8 abgr_to_y_1920_avx: 926.8 abgr_to_y_1920_avx2: 519.3 bgra_to_uv_8_c: 37.3 bgra_to_uv_8_sse2: 13.3 bgra_to_uv_8_avx: 14.8 bgra_to_uv_8_avx2: 19.8 bgra_to_uv_128_c: 563.8 bgra_to_uv_128_sse2: 111.3 bgra_to_uv_128_avx: 112.3 bgra_to_uv_128_avx2: 64.8 bgra_to_uv_1080_c: 4691.8 bgra_to_uv_1080_sse2: 893.8 bgra_to_uv_1080_avx: 899.8 bgra_to_uv_1080_avx2: 517.8 bgra_to_uv_1920_c: 8332.8 bgra_to_uv_1920_sse2: 1590.8 bgra_to_uv_1920_avx: 1605.8 bgra_to_uv_1920_avx2: 867.3 bgra_to_y_8_c: 22.3 bgra_to_y_8_sse2: 12.8 bgra_to_y_8_avx: 12.8 bgra_to_y_8_avx2: 17.3 bgra_to_y_128_c: 291.3 bgra_to_y_128_sse2: 67.8 bgra_to_y_128_avx: 69.3 bgra_to_y_128_avx2: 45.3 bgra_to_y_1080_c: 2357.3 bgra_to_y_1080_sse2: 508.3 bgra_to_y_1080_avx: 518.3 bgra_to_y_1080_avx2: 399.8 bgra_to_y_1920_c: 4202.8 bgra_to_y_1920_sse2: 906.8 bgra_to_y_1920_avx: 907.3 bgra_to_y_1920_avx2: 526.3 Signed-off-by: James Almer --- libswscale/x86/input.asm | 51 ++++++++++++++++++++++++++++++++++++---- libswscale/x86/swscale.c | 8 +++++++ 2 files changed, 55 insertions(+), 4 deletions(-) diff --git a/libswscale/x86/input.asm b/libswscale/x86/input.asm index e79fe11405..f1ad6a53fd 100644 --- a/libswscale/x86/input.asm +++ b/libswscale/x86/input.asm @@ -380,8 +380,13 @@ RGB24_FUNCS 11, 13 ; %2-5 = rgba, bgra, argb or abgr (in individual characters) %macro RGB32_TO_Y_FN 5-6 cglobal %2%3%4%5 %+ ToY, 6, 6, %1, dst, src, u1, u2, w, table +%if mmsize == 32 + vbroadcasti128 m5, [rgba_Ycoeff_%2%4] + vbroadcasti128 m6, [rgba_Ycoeff_%3%5] +%else mova m5, [rgba_Ycoeff_%2%4] mova m6, [rgba_Ycoeff_%3%5] +%endif %if %0 == 6 jmp mangle(private_prefix %+ _ %+ %6 %+ ToY %+ SUFFIX).body %else ; %0 == 6 @@ -394,13 +399,21 @@ cglobal %2%3%4%5 %+ ToY, 6, 6, %1, dst, src, u1, u2, w, table lea srcq, [srcq+wq*2] add dstq, wq neg wq +%if mmsize == 32 + vbroadcasti128 m4, [rgb_Yrnd] +%else mova m4, [rgb_Yrnd] +%endif pcmpeqb m7, m7 psrlw m7, 8 ; (word) { 0x00ff } x4 .loop: ; FIXME check alignment and use mova - movu m0, [srcq+wq*2+0] ; (byte) { Bx, Gx, Rx, xx }[0-3] - movu m2, [srcq+wq*2+mmsize] ; (byte) { Bx, Gx, Rx, xx }[4-7] + movu xm0, [srcq+wq*2+0] ; (byte) { Bx, Gx, Rx, xx }[0-3] + movu xm2, [srcq+wq*2+16] ; (byte) { Bx, Gx, Rx, xx }[4-7] +%if mmsize == 32 + vinserti128 m0, m0, [srcq+wq*2+32], 1 + vinserti128 m2, m2, [srcq+wq*2+48], 1 +%endif DEINTB 1, 0, 3, 2, 7 ; (word) { Gx, xx (m0/m2) or Bx, Rx (m1/m3) }[0-3]/[4-7] pmaddwd m1, m5 ; (dword) { Bx*BY + Rx*RY }[0-3] pmaddwd m0, m6 ; (dword) { Gx*GY }[0-3] @@ -421,6 +434,7 @@ cglobal %2%3%4%5 %+ ToY, 6, 6, %1, dst, src, u1, u2, w, table add srcq, 2*mmsize - 2 add dstq, mmsize - 1 .loop2: +INIT_XMM cpuname movd m0, [srcq+wq*2+0] ; (byte) { Bx, Gx, Rx, xx }[0-3] DEINTB 1, 0, 3, 2, 7 ; (word) { Gx, xx (m0/m2) or Bx, Rx (m1/m3) }[0-3]/[4-7] pmaddwd m1, m5 ; (dword) { Bx*BY + Rx*RY }[0-3] @@ -433,6 +447,9 @@ cglobal %2%3%4%5 %+ ToY, 6, 6, %1, dst, src, u1, u2, w, table add wq, 2 jl .loop2 .end: +%if cpuflag(avx2) +INIT_YMM cpuname +%endif RET %endif ; %0 == 3 %endmacro @@ -442,10 +459,17 @@ cglobal %2%3%4%5 %+ ToY, 6, 6, %1, dst, src, u1, u2, w, table %macro RGB32_TO_UV_FN 5-6 cglobal %2%3%4%5 %+ ToUV, 7, 7, %1, dstU, dstV, u1, src, u2, w, table %if ARCH_X86_64 +%if mmsize == 32 + vbroadcasti128 m8, [rgba_Ucoeff_%2%4] + vbroadcasti128 m9, [rgba_Ucoeff_%3%5] + vbroadcasti128 m10, [rgba_Vcoeff_%2%4] + vbroadcasti128 m11, [rgba_Vcoeff_%3%5] +%else mova m8, [rgba_Ucoeff_%2%4] mova m9, [rgba_Ucoeff_%3%5] mova m10, [rgba_Vcoeff_%2%4] mova m11, [rgba_Vcoeff_%3%5] +%endif %define coeffU1 m8 %define coeffU2 m9 %define coeffV1 m10 @@ -473,11 +497,19 @@ cglobal %2%3%4%5 %+ ToUV, 7, 7, %1, dstU, dstV, u1, src, u2, w, table neg wq pcmpeqb m7, m7 psrlw m7, 8 ; (word) { 0x00ff } x4 +%if mmsize == 32 + vbroadcasti128 m6, [rgb_UVrnd] +%else mova m6, [rgb_UVrnd] +%endif .loop: ; FIXME check alignment and use mova - movu m0, [srcq+wq*2+0] ; (byte) { Bx, Gx, Rx, xx }[0-3] - movu m4, [srcq+wq*2+mmsize] ; (byte) { Bx, Gx, Rx, xx }[4-7] + movu xm0, [srcq+wq*2+0] ; (byte) { Bx, Gx, Rx, xx }[0-3] + movu xm4, [srcq+wq*2+16] ; (byte) { Bx, Gx, Rx, xx }[4-7] +%if mmsize == 32 + vinserti128 m0, m0, [srcq+wq*2+32], 1 + vinserti128 m4, m4, [srcq+wq*2+48], 1 +%endif DEINTB 1, 0, 5, 4, 7 ; (word) { Gx, xx (m0/m4) or Bx, Rx (m1/m5) }[0-3]/[4-7] pmaddwd m3, m1, coeffV1 ; (dword) { Bx*BV + Rx*RV }[0-3] pmaddwd m2, m0, coeffV2 ; (dword) { Gx*GV }[0-3] @@ -511,6 +543,7 @@ cglobal %2%3%4%5 %+ ToUV, 7, 7, %1, dstU, dstV, u1, src, u2, w, table add dstUq, mmsize - 1 add dstVq, mmsize - 1 .loop2: +INIT_XMM cpuname movd m0, [srcq+wq*2] ; (byte) { Bx, Gx, Rx, xx }[0-3] DEINTB 1, 0, 5, 4, 7 ; (word) { Gx, xx (m0/m4) or Bx, Rx (m1/m5) }[0-3]/[4-7] pmaddwd m3, m1, coeffV1 ; (dword) { Bx*BV + Rx*RV }[0-3] @@ -530,6 +563,9 @@ cglobal %2%3%4%5 %+ ToUV, 7, 7, %1, dstU, dstV, u1, src, u2, w, table add wq, 2 jl .loop2 .end: +%if cpuflag(avx2) +INIT_YMM cpuname +%endif RET %endif ; ARCH_X86_64 && %0 == 3 %endmacro @@ -556,6 +592,13 @@ INIT_XMM avx RGB32_FUNCS 8, 12 %endif +%if ARCH_X86_64 +%if HAVE_AVX2_EXTERNAL +INIT_YMM avx2 +RGB32_FUNCS 8, 12 +%endif +%endif + ;----------------------------------------------------------------------------- ; YUYV/UYVY/NV12/NV21 packed pixel shuffling. ; diff --git a/libswscale/x86/swscale.c b/libswscale/x86/swscale.c index 1438c077e6..5a9da23265 100644 --- a/libswscale/x86/swscale.c +++ b/libswscale/x86/swscale.c @@ -321,6 +321,10 @@ void ff_ ## fmt ## ToUV_ ## opt(uint8_t *dstU, uint8_t *dstV, \ INPUT_FUNCS(sse2); INPUT_FUNCS(ssse3); INPUT_FUNCS(avx); +INPUT_FUNC(rgba, avx2); +INPUT_FUNC(bgra, avx2); +INPUT_FUNC(argb, avx2); +INPUT_FUNC(abgr, avx2); INPUT_FUNC(rgb24, avx2); INPUT_FUNC(bgr24, avx2); @@ -640,6 +644,10 @@ switch(c->dstBpc){ \ switch (c->srcFormat) { case_rgb(rgb24, RGB24, avx2); case_rgb(bgr24, BGR24, avx2); + case_rgb(bgra, BGRA, avx2); + case_rgb(rgba, RGBA, avx2); + case_rgb(abgr, ABGR, avx2); + case_rgb(argb, ARGB, avx2); } switch (c->dstFormat) { case AV_PIX_FMT_NV12: From patchwork Wed Jun 5 20:28:53 2024 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: James Almer X-Patchwork-Id: 49604 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a59:9185:0:b0:460:55fa:d5ed with SMTP id s5csp15805vqg; Wed, 5 Jun 2024 13:39:08 -0700 (PDT) X-Forwarded-Encrypted: i=2; AJvYcCX88ozl+U64nCo0Y69ElMYW3yGt10zxuDTox0q8aNHhajTiTUlic1ykMZ41SnkPwr6nYeLloHtsOcOaJW3MqDVClNGJsREtAw6BwQ== X-Google-Smtp-Source: AGHT+IGKWnKfw1bXppPeiHYZB6MA2CLsyf7QYMf9idGxAa0fiHlUeKbwxYLiXQhWqfPC5cBYZr2W X-Received: by 2002:a17:907:38c:b0:a68:a800:5f7e with SMTP id a640c23a62f3a-a699faa9160mr239260466b.10.1717619947913; Wed, 05 Jun 2024 13:39:07 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1717619947; cv=none; d=google.com; s=arc-20160816; b=TJNSKfFtqhf8rK04x101LMBwurpjL3ZJjIMdLhtsIEzjy6lscIiFN0lGKoeDH4SSM3 Q47I1EdDFkMKBR0u4J8PBGIt3pYMwriRnUxRdI3rf7DynVNhPSZxMfHrMNW87+NVq9uG rP58ZGgAhtgBc+5vJCPzKxK7jVOLLSp62C0qG//PaKraDmqfBlU0svY7oNe1A7fj5qVe JnAV/DHVwX7qcf9t4ml5+AKVshTF/4wX2Aed3ZbHibIuj87m495xrhPoraAZyVNm/Tzj /atmaWVCxOKlytabN9k7ChlMmWqNybB727/cHY1fUy8Rh5xSGoeqW9IE4K/ta7qjpDLX lmvw== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:reply-to:list-subscribe :list-help:list-post:list-archive:list-unsubscribe:list-id :precedence:subject:mime-version:references:in-reply-to:message-id :date:to:from:dkim-signature:delivered-to; bh=uQrBz6pgPxpE6oWIF79jFwN4dg80e2ztCAnCa6P0S/w=; fh=YOA8vD9MJZuwZ71F/05pj6KdCjf6jQRmzLS+CATXUQk=; b=zvwp+EvMJzDqWI/WIf4q1EpI/qWUyij1osNsCmBsRtHxOwCgJgENNwHF9Jq7JfddGU kTE00ax0Pb4znlnacKfcZ7M2GBXtCsbk/Mzcplci13YkkWyZ42KSXEYRvVWRF7v3Sz/j G9SPHOet2evzR394C35TdO/cj3c19l2SXixCi42zAGeKj93pJ4z4rlpSiV/9YtRHsxRS 8fepU9iR578VZOhyb1qUXR/zoEB23cKOr8oO7DmXL5QKX5q9A0gcoO7kVM7o2UpT69nF KB25FyyJChwZ/s4nHi3fNaVjDEdCWbCn464lKoS5VNwkGuB/Bd42kiq4vPhg9u2I5EW0 28Vg==; dara=google.com ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@gmail.com header.s=20230601 header.b=W17YWjwA; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=QUARANTINE dis=NONE) header.from=gmail.com Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id a640c23a62f3a-a68e37c13b7si466781466b.352.2024.06.05.13.39.06; Wed, 05 Jun 2024 13:39:07 -0700 (PDT) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@gmail.com header.s=20230601 header.b=W17YWjwA; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=QUARANTINE dis=NONE) header.from=gmail.com Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 4DDE068D44A; Wed, 5 Jun 2024 23:29:20 +0300 (EEST) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-pf1-f178.google.com (mail-pf1-f178.google.com [209.85.210.178]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 4CA1868D44A for ; Wed, 5 Jun 2024 23:29:14 +0300 (EEST) Received: by mail-pf1-f178.google.com with SMTP id d2e1a72fcca58-70249faa853so155135b3a.3 for ; Wed, 05 Jun 2024 13:29:14 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1717619351; x=1718224151; darn=ffmpeg.org; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:to:from:from:to:cc:subject:date:message-id :reply-to; bh=P0lzJOk91F7ggDylIlcAozW5VSVZjok2hpme3e8CLpk=; b=W17YWjwAmUl0BJTAAZCSrsjSRmio/bMZcOYBmrGFtYNbYNFVHIs55cZy4Zk6ls5T2u SeRGVIbiAPlSM6Av/NMvdHkEOSl4CV4sXhKCtlyWifkBFGfrDdLheXmbix+22B6vpFcK KQecMxSeFiLmcUZlga8s1PMDpqIw+az1svR9lscifm+PGjCnA5URjG/dkrwN9gE1SQtT kZZZTMXWs0IfWWlbHFRpUWltTqmnWEQX8ckaDBEZY8+QDTd/GoRQ60lLk8CMbK65krBS RBrymBYvBNZ7Y7ayuM3GA4a9N5y0Zs90BLdLwKUSpHRvWLlLtafnSwAGDVuB+PGNxfAA Uw5A== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1717619351; x=1718224151; h=content-transfer-encoding:mime-version:references:in-reply-to :message-id:date:subject:to:from:x-gm-message-state:from:to:cc :subject:date:message-id:reply-to; bh=P0lzJOk91F7ggDylIlcAozW5VSVZjok2hpme3e8CLpk=; b=F1HRNzMHy+9R1Wj1cXlTCkYK9VOVg6nIeODVBOi6QM7Ru1FWHVKRPRXL3osDxWEVVM 3z2IcaZllAUlnekuesOvQhexvFsQxdN4qaLw/kR5SSovZUC8ZufSbzcyMDa7UKeTeRMY yHpU3nXEpUHsZr4djMFl7L1JssN388QOaWgA+DJOVF/wlnk50Zb+3Y3fCug5lGz2LFLC lrVyIKEtPLpf0GRzaukwKoxBTMsIa4PIAPvZtgc7epP0o496yoZUP2QXQCIjBp1YhJbW VkFfiykjJFM9bCzXnoc+c8t5yU5+SqnRT80yFKe45tS1aER/snv8dmsIiGx1WwKNcTab k2pw== X-Gm-Message-State: AOJu0YzXxRlRqfSnU9Qz6iXl32dPhFt9W+UgVDAx/RJZqMbplBwXPr8o 8AjQDs1vEWzvHGYoOivP/zflPveqKmI5dPHdHugI+pK9h+o8fOkme64sDA== X-Received: by 2002:a05:6a20:918e:b0:1af:cbd3:ab4c with SMTP id adf61e73a8af0-1b2b7025e81mr4665052637.35.1717619351341; Wed, 05 Jun 2024 13:29:11 -0700 (PDT) Received: from localhost.localdomain ([190.194.167.233]) by smtp.gmail.com with ESMTPSA id d2e1a72fcca58-7024967ae05sm8536492b3a.157.2024.06.05.13.29.10 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Wed, 05 Jun 2024 13:29:10 -0700 (PDT) From: James Almer To: ffmpeg-devel@ffmpeg.org Date: Wed, 5 Jun 2024 17:28:53 -0300 Message-ID: <20240605202853.3135-2-jamrial@gmail.com> X-Mailer: git-send-email 2.45.1 In-Reply-To: <20240605202853.3135-1-jamrial@gmail.com> References: <20240605202853.3135-1-jamrial@gmail.com> MIME-Version: 1.0 Subject: [FFmpeg-devel] [PATCH 2/2] swscale/x86/input: add AVX2 optimized uyvytoyuv422 X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: WF6gVAbAuW+9 uyvytoyuv422_c: 23991.8 uyvytoyuv422_sse2: 2817.8 uyvytoyuv422_avx: 2819.3 uyvytoyuv422_avx2: 1972.3 Signed-off-by: James Almer --- libswscale/x86/rgb2rgb.c | 6 ++++++ libswscale/x86/rgb_2_rgb.asm | 32 ++++++++++++++++++++++++-------- 2 files changed, 30 insertions(+), 8 deletions(-) diff --git a/libswscale/x86/rgb2rgb.c b/libswscale/x86/rgb2rgb.c index b325e5dbd5..21ccfafe51 100644 --- a/libswscale/x86/rgb2rgb.c +++ b/libswscale/x86/rgb2rgb.c @@ -136,6 +136,9 @@ void ff_uyvytoyuv422_sse2(uint8_t *ydst, uint8_t *udst, uint8_t *vdst, void ff_uyvytoyuv422_avx(uint8_t *ydst, uint8_t *udst, uint8_t *vdst, const uint8_t *src, int width, int height, int lumStride, int chromStride, int srcStride); +void ff_uyvytoyuv422_avx2(uint8_t *ydst, uint8_t *udst, uint8_t *vdst, + const uint8_t *src, int width, int height, + int lumStride, int chromStride, int srcStride); #endif av_cold void rgb2rgb_init_x86(void) @@ -177,5 +180,8 @@ av_cold void rgb2rgb_init_x86(void) if (EXTERNAL_AVX(cpu_flags)) { uyvytoyuv422 = ff_uyvytoyuv422_avx; } + if (EXTERNAL_AVX2_FAST(cpu_flags)) { + uyvytoyuv422 = ff_uyvytoyuv422_avx2; + } #endif } diff --git a/libswscale/x86/rgb_2_rgb.asm b/libswscale/x86/rgb_2_rgb.asm index 76ca1eec03..0bf1278718 100644 --- a/libswscale/x86/rgb_2_rgb.asm +++ b/libswscale/x86/rgb_2_rgb.asm @@ -34,13 +34,16 @@ pb_shuffle3210: db 3, 2, 1, 0, 7, 6, 5, 4, 11, 10, 9, 8, 15, 14, 13, 12 SECTION .text -%macro RSHIFT_COPY 3 +%macro RSHIFT_COPY 5 ; %1 dst ; %2 src ; %3 shift -%if cpuflag(avx) - psrldq %1, %2, %3 +%if mmsize == 32 + vperm2i128 %1, %2, %3, %5 + RSHIFT %1, %4 +%elif cpuflag(avx) + psrldq %1, %2, %4 %else mova %1, %2 - RSHIFT %1, %3 + RSHIFT %1, %4 %endif %endmacro @@ -233,26 +236,37 @@ cglobal uyvytoyuv422, 9, 14, 8, ydst, udst, vdst, src, w, h, lum_stride, chrom_s jge .end_line .loop_simd: +%if mmsize == 32 + movu xm2, [srcq + wtwoq ] + movu xm3, [srcq + wtwoq + 16 ] + movu xm4, [srcq + wtwoq + 16 * 2] + movu xm5, [srcq + wtwoq + 16 * 3] + vinserti128 m2, m2, [srcq + wtwoq + 16 * 4], 1 + vinserti128 m3, m3, [srcq + wtwoq + 16 * 5], 1 + vinserti128 m4, m4, [srcq + wtwoq + 16 * 6], 1 + vinserti128 m5, m5, [srcq + wtwoq + 16 * 7], 1 +%else movu m2, [srcq + wtwoq ] movu m3, [srcq + wtwoq + mmsize ] movu m4, [srcq + wtwoq + mmsize * 2] movu m5, [srcq + wtwoq + mmsize * 3] +%endif ; extract y part 1 - RSHIFT_COPY m6, m2, 1 ; UYVY UYVY -> YVYU YVY... + RSHIFT_COPY m6, m2, m4, 1, 0x20 ; UYVY UYVY -> YVYU YVY... pand m6, m1; YxYx YxYx... - RSHIFT_COPY m7, m3, 1 ; UYVY UYVY -> YVYU YVY... + RSHIFT_COPY m7, m3, m5, 1, 0x20 ; UYVY UYVY -> YVYU YVY... pand m7, m1 ; YxYx YxYx... packuswb m6, m7 ; YYYY YYYY... movu [ydstq + wq], m6 ; extract y part 2 - RSHIFT_COPY m6, m4, 1 ; UYVY UYVY -> YVYU YVY... + RSHIFT_COPY m6, m4, m2, 1, 0x13 ; UYVY UYVY -> YVYU YVY... pand m6, m1; YxYx YxYx... - RSHIFT_COPY m7, m5, 1 ; UYVY UYVY -> YVYU YVY... + RSHIFT_COPY m7, m5, m3, 1, 0x13 ; UYVY UYVY -> YVYU YVY... pand m7, m1 ; YxYx YxYx... packuswb m6, m7 ; YYYY YYYY... @@ -309,4 +323,6 @@ UYVY_TO_YUV422 INIT_XMM avx UYVY_TO_YUV422 +INIT_YMM avx2 +UYVY_TO_YUV422 %endif