From patchwork Tue Jun 4 19:07:42 2024 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: James Almer X-Patchwork-Id: 49560 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a59:d153:0:b0:460:55fa:d5ed with SMTP id bt19csp101897vqb; Tue, 4 Jun 2024 12:08:10 -0700 (PDT) X-Forwarded-Encrypted: i=2; AJvYcCWyfy0YmwssCHs/QCoh7B/Wvw1JXD42BqcCpQb27Cp53XpA/dae9rEpJCVKK1YTd10QbvUhLMQGCiC4Zs/o77WCs6wLvXwYN2DNYg== X-Google-Smtp-Source: AGHT+IFt6PqhoxSrrMCMh34+BNKqOCC+zRpw+KPyMg6pmXgARICnr98nTyspO86q5mNXnz1LvqyG X-Received: by 2002:a17:906:c198:b0:a63:582b:8ac1 with SMTP id a640c23a62f3a-a699face692mr25351566b.20.1717528089837; Tue, 04 Jun 2024 12:08:09 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1717528089; cv=none; d=google.com; s=arc-20160816; b=SddreLVNpAZta7R0evdcfWrWrTI53ARQY4yJJtCYXpeUcEf29WbeqydHZrShu/z4e4 C0LCIcOAlmv0uFm7cYs7u85O3tEGItZtXur56hMl0+j5Cy2lj84bif6lALLzq1WEAm2P 82j8oATiOXBgUiyBLvMEQXW8jmMrn1siRiOT4RYJ30iiluuGXt/T/lnstMtGpQxcxKhN Sbmli1nv1P/cCsv6Z9QtW0MrJXH5N0Ru8FSseqKkNbBwx3GVHucG6uAYUKIDfGix4a5N NaW07xp3LdaO/OXWmBTXz8fBsYmjveShrXE2OMcUe5aNkmpSm5CP5dcJZt4MM7LETf+I 0AYA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:reply-to:list-subscribe :list-help:list-post:list-archive:list-unsubscribe:list-id :precedence:subject:mime-version:message-id:date:to:from :dkim-signature:delivered-to; bh=+QHfeRAgwlDWSBsJMuhBjChrx03J2sBGWh2ZsHMySoE=; fh=YOA8vD9MJZuwZ71F/05pj6KdCjf6jQRmzLS+CATXUQk=; b=DDLY6eXxj5KldOaiGLLU4LvkBSKZTSggBVKU3Iq53jq0ce5lHzYrTD26ym5VBoYtw/ ZkfeezM7jiRTiMOLMYHLhdoIwImqyDsherHTf7uunaUHZg+QOag+yJ5U6lJtdSyazw+E x6NW8CK9SHlshDzAqdv8c//ICH97BBqt+8o3bjPj0m3N9+5ujvXwx9d/foltISMjI5z6 is9536C+jSEDkO4OIdQ5896xmPLfXphdVlt5PomHrYW9cCCxcnAcw5w8Yq8qSQhBJYhr tAYpQ1aqMYBQ52KmUI6vcocCwOJ7rCdpvIZcmmNFvuFLkgVABv/dnw35BfN5Za7F0UmM W+tQ==; dara=google.com ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@gmail.com header.s=20230601 header.b=MdNSfKDz; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=QUARANTINE dis=NONE) header.from=gmail.com Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id a640c23a62f3a-a68f117bf7dsi300113266b.870.2024.06.04.12.08.09; Tue, 04 Jun 2024 12:08:09 -0700 (PDT) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@gmail.com header.s=20230601 header.b=MdNSfKDz; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=QUARANTINE dis=NONE) header.from=gmail.com Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 74A1E68D462; Tue, 4 Jun 2024 22:08:06 +0300 (EEST) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-pg1-f179.google.com (mail-pg1-f179.google.com [209.85.215.179]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id EBB1268D462 for ; Tue, 4 Jun 2024 22:07:58 +0300 (EEST) Received: by mail-pg1-f179.google.com with SMTP id 41be03b00d2f7-6c53a315c6eso3447078a12.3 for ; Tue, 04 Jun 2024 12:07:58 -0700 (PDT) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=gmail.com; s=20230601; t=1717528076; x=1718132876; darn=ffmpeg.org; h=content-transfer-encoding:mime-version:message-id:date:subject:to :from:from:to:cc:subject:date:message-id:reply-to; bh=3CYxBbWIrgLG6uU0peVbCSBR6cODJto/FRyeobvYUzE=; b=MdNSfKDzAf/OFiguaYy3kVwEwM02mYme2niULrrCIf6kkIdm+X+lwCIYb4XpEuXdxq 2GgSlOtvFxP8G0lpqbMuQynke8XCK8FYecDFtjuEpEEB9w9XKVW8/mYQc5wq+IWuLk6Q nxHotAh6zqslNJl86VMlOn5kP6fpSF/MpQXGm1HcZ23TvFiYP0xkdzZADUIIlNnkP1OV Ul+c7KmaNFXlLX4V3Jm2Si2dROf/xPOtHbyue0JQQKQFGyVGgLqoQhiIskr2jR2rb5hH Mc/8I+Z2o6a4bVMgg7Mcg7M2k/+lprY2c9KHliKenKLsYIEuQPvTxEpt4zFF5j5LaCru CkBw== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20230601; t=1717528076; x=1718132876; h=content-transfer-encoding:mime-version:message-id:date:subject:to :from:x-gm-message-state:from:to:cc:subject:date:message-id:reply-to; bh=3CYxBbWIrgLG6uU0peVbCSBR6cODJto/FRyeobvYUzE=; b=wmd8ejAZIASv0eFz6eLWq5+RHCSJRNGH5A0Oa+nnt9AWDUU0G9DHRKqLlVwrbXoMaZ wFLdBasxUWb/W7hHJopWhbOlfDyFZInTaGSAE/bUt//gWhWw1CH3Nrob/spCEYIg5vY2 UIakeMNYpGsDPKqSkmlZtWpLuxiBA7IKbTI4bz+kKMdzPz7g8hT/3WTC8AKHSZHTEX35 C+68foLrbXPPqRMo+uFCB5CcEXcSDBPh654MiDk2i4nc6WOeeE4h/2qHMCR/PqWq1ljR KrGDTpxXkQP02ED8vkgJEZH6SQN034r4ETxdpj0fqzOwMPHmmuris47zLs4tBiZx6GeU iGjg== X-Gm-Message-State: AOJu0Yy0Z1l8iDAXsydj0U08CCG/1OP8KQyIycBm+UVUlzUqzI5Th0cK A9+J/+ibxmmEAJJrPx91xXQdI3N+lI4KT/z1DvYlAQ7JlXNnd/oFViWlzg== X-Received: by 2002:a17:90b:110f:b0:2bd:d6cc:c305 with SMTP id 98e67ed59e1d1-2c27dd578bfmr321806a91.49.1717528076369; Tue, 04 Jun 2024 12:07:56 -0700 (PDT) Received: from localhost.localdomain ([190.194.167.233]) by smtp.gmail.com with ESMTPSA id 98e67ed59e1d1-2c1a77bb65bsm10197120a91.56.2024.06.04.12.07.54 for (version=TLS1_3 cipher=TLS_AES_256_GCM_SHA384 bits=256/256); Tue, 04 Jun 2024 12:07:55 -0700 (PDT) From: James Almer To: ffmpeg-devel@ffmpeg.org Date: Tue, 4 Jun 2024 16:07:42 -0300 Message-ID: <20240604190742.1742-1-jamrial@gmail.com> X-Mailer: git-send-email 2.45.1 MIME-Version: 1.0 Subject: [FFmpeg-devel] [PATCH] swscale/x86/input: add AVX2 optimized RGB24 to YUV functions X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: 39sNPOgOYJ+q rgb24_to_uv_8_c: 39.3 rgb24_to_uv_8_sse2: 14.3 rgb24_to_uv_8_ssse3: 13.3 rgb24_to_uv_8_avx: 12.8 rgb24_to_uv_8_avx2: 14.3 rgb24_to_uv_128_c: 582.8 rgb24_to_uv_128_sse2: 127.3 rgb24_to_uv_128_ssse3: 107.3 rgb24_to_uv_128_avx: 111.3 rgb24_to_uv_128_avx2: 62.3 rgb24_to_uv_1080_c: 4981.3 rgb24_to_uv_1080_sse2: 1048.3 rgb24_to_uv_1080_ssse3: 876.8 rgb24_to_uv_1080_avx: 887.8 rgb24_to_uv_1080_avx2: 492.3 rgb24_to_uv_1280_c: 5906.8 rgb24_to_uv_1280_sse2: 1263.3 rgb24_to_uv_1280_ssse3: 1048.3 rgb24_to_uv_1280_avx: 1045.8 rgb24_to_uv_1280_avx2: 579.8 rgb24_to_uv_1920_c: 8665.3 rgb24_to_uv_1920_sse2: 1888.8 rgb24_to_uv_1920_ssse3: 1571.8 rgb24_to_uv_1920_avx: 1558.8 rgb24_to_uv_1920_avx2: 869.3 rgb24_to_y_8_c: 20.3 rgb24_to_y_8_sse2: 11.8 rgb24_to_y_8_ssse3: 10.3 rgb24_to_y_8_avx: 10.3 rgb24_to_y_8_avx2: 10.8 rgb24_to_y_128_c: 284.8 rgb24_to_y_128_sse2: 83.3 rgb24_to_y_128_ssse3: 66.8 rgb24_to_y_128_avx: 64.8 rgb24_to_y_128_avx2: 39.3 rgb24_to_y_1080_c: 2451.3 rgb24_to_y_1080_sse2: 696.3 rgb24_to_y_1080_ssse3: 516.8 rgb24_to_y_1080_avx: 518.8 rgb24_to_y_1080_avx2: 301.8 rgb24_to_y_1280_c: 2892.8 rgb24_to_y_1280_sse2: 816.8 rgb24_to_y_1280_ssse3: 623.3 rgb24_to_y_1280_avx: 616.3 rgb24_to_y_1280_avx2: 350.8 rgb24_to_y_1920_c: 4338.8 rgb24_to_y_1920_sse2: 1210.8 rgb24_to_y_1920_ssse3: 928.3 rgb24_to_y_1920_avx: 920.3 rgb24_to_y_1920_avx2: 534.8 Signed-off-by: James Almer --- libswscale/x86/input.asm | 49 ++++++++++++++++++++++++++++++++++++---- libswscale/x86/swscale.c | 7 ++++++ 2 files changed, 51 insertions(+), 5 deletions(-) diff --git a/libswscale/x86/input.asm b/libswscale/x86/input.asm index a197183f1f..e79fe11405 100644 --- a/libswscale/x86/input.asm +++ b/libswscale/x86/input.asm @@ -23,7 +23,7 @@ %include "libavutil/x86/x86util.asm" -SECTION_RODATA +SECTION_RODATA 32 %define RY 0x20DE %define GY 0x4087 @@ -90,8 +90,12 @@ rgb_UVrnd: times 4 dd 0x400100 ; 128.5 << 15 ; rgba_Vcoeff_ag: times 4 dw 0, GV shuf_rgb_12x4: db 0, 0x80, 1, 0x80, 2, 0x80, 3, 0x80, \ + 6, 0x80, 7, 0x80, 8, 0x80, 9, 0x80, \ + 0, 0x80, 1, 0x80, 2, 0x80, 3, 0x80, \ 6, 0x80, 7, 0x80, 8, 0x80, 9, 0x80 shuf_rgb_3x56: db 2, 0x80, 3, 0x80, 4, 0x80, 5, 0x80, \ + 8, 0x80, 9, 0x80, 10, 0x80, 11, 0x80, \ + 2, 0x80, 3, 0x80, 4, 0x80, 5, 0x80, \ 8, 0x80, 9, 0x80, 10, 0x80, 11, 0x80 pd_65535f: times 8 dd 65535.0 pb_pack_shuffle16le: db 0, 1, 4, 5, \ @@ -134,8 +138,13 @@ SECTION .text %macro RGB24_TO_Y_FN 2-3 cglobal %2 %+ 24ToY, 6, 6, %1, dst, src, u1, u2, w, table %if ARCH_X86_64 +%if mmsize == 32 + vbroadcasti128 m8, [%2_Ycoeff_12x4] + vbroadcasti128 m9, [%2_Ycoeff_3x56] +%else mova m8, [%2_Ycoeff_12x4] mova m9, [%2_Ycoeff_3x56] +%endif %define coeff1 m8 %define coeff2 m9 %else ; x86-32 @@ -165,11 +174,19 @@ cglobal %2 %+ 24ToY, 6, 6, %1, dst, src, u1, u2, w, table %if notcpuflag(ssse3) pxor m7, m7 %endif ; !cpuflag(ssse3) +%if mmsize == 32 + vbroadcasti128 m4, [rgb_Yrnd] +%else mova m4, [rgb_Yrnd] +%endif .loop: %if cpuflag(ssse3) - movu m0, [srcq+0] ; (byte) { Bx, Gx, Rx }[0-3] - movu m2, [srcq+12] ; (byte) { Bx, Gx, Rx }[4-7] + movu xm0, [srcq+0] ; (byte) { Bx, Gx, Rx }[0-3] + movu xm2, [srcq+12] ; (byte) { Bx, Gx, Rx }[4-7] +%if mmsize == 32 + vinserti128 m0, m0, [srcq+24], 1 + vinserti128 m2, m2, [srcq+36], 1 +%endif pshufb m1, m0, shuf_rgb2 ; (word) { R0, B1, G1, R1, R2, B3, G3, R3 } pshufb m0, shuf_rgb1 ; (word) { B0, G0, R0, B1, B2, G2, R2, B3 } pshufb m3, m2, shuf_rgb2 ; (word) { R4, B5, G5, R5, R6, B7, G7, R7 } @@ -216,10 +233,17 @@ cglobal %2 %+ 24ToY, 6, 6, %1, dst, src, u1, u2, w, table %macro RGB24_TO_UV_FN 2-3 cglobal %2 %+ 24ToUV, 7, 7, %1, dstU, dstV, u1, src, u2, w, table %if ARCH_X86_64 +%if mmsize == 32 + vbroadcasti128 m8, [%2_Ucoeff_12x4] + vbroadcasti128 m9, [%2_Ucoeff_3x56] + vbroadcasti128 m10, [%2_Vcoeff_12x4] + vbroadcasti128 m11, [%2_Vcoeff_3x56] +%else mova m8, [%2_Ucoeff_12x4] mova m9, [%2_Ucoeff_3x56] mova m10, [%2_Vcoeff_12x4] mova m11, [%2_Vcoeff_3x56] +%endif %define coeffU1 m8 %define coeffU2 m9 %define coeffV1 m10 @@ -253,14 +277,22 @@ cglobal %2 %+ 24ToUV, 7, 7, %1, dstU, dstV, u1, src, u2, w, table add dstUq, wq add dstVq, wq neg wq +%if mmsize == 32 + vbroadcasti128 m6, [rgb_UVrnd] +%else mova m6, [rgb_UVrnd] +%endif %if notcpuflag(ssse3) pxor m7, m7 %endif .loop: %if cpuflag(ssse3) - movu m0, [srcq+0] ; (byte) { Bx, Gx, Rx }[0-3] - movu m4, [srcq+12] ; (byte) { Bx, Gx, Rx }[4-7] + movu xm0, [srcq+0] ; (byte) { Bx, Gx, Rx }[0-3] + movu xm4, [srcq+12] ; (byte) { Bx, Gx, Rx }[4-7] +%if mmsize == 32 + vinserti128 m0, m0, [srcq+24], 1 + vinserti128 m4, m4, [srcq+36], 1 +%endif pshufb m1, m0, shuf_rgb2 ; (word) { R0, B1, G1, R1, R2, B3, G3, R3 } pshufb m0, shuf_rgb1 ; (word) { B0, G0, R0, B1, B2, G2, R2, B3 } %else ; !cpuflag(ssse3) @@ -337,6 +369,13 @@ INIT_XMM avx RGB24_FUNCS 11, 13 %endif +%if ARCH_X86_64 +%if HAVE_AVX2_EXTERNAL +INIT_YMM avx2 +RGB24_FUNCS 11, 13 +%endif +%endif + ; %1 = nr. of XMM registers ; %2-5 = rgba, bgra, argb or abgr (in individual characters) %macro RGB32_TO_Y_FN 5-6 diff --git a/libswscale/x86/swscale.c b/libswscale/x86/swscale.c index fff8bb4396..1438c077e6 100644 --- a/libswscale/x86/swscale.c +++ b/libswscale/x86/swscale.c @@ -321,6 +321,8 @@ void ff_ ## fmt ## ToUV_ ## opt(uint8_t *dstU, uint8_t *dstV, \ INPUT_FUNCS(sse2); INPUT_FUNCS(ssse3); INPUT_FUNCS(avx); +INPUT_FUNC(rgb24, avx2); +INPUT_FUNC(bgr24, avx2); #if ARCH_X86_64 #define YUV2NV_DECL(fmt, opt) \ @@ -634,6 +636,11 @@ switch(c->dstBpc){ \ } if (EXTERNAL_AVX2_FAST(cpu_flags)) { + if (ARCH_X86_64) + switch (c->srcFormat) { + case_rgb(rgb24, RGB24, avx2); + case_rgb(bgr24, BGR24, avx2); + } switch (c->dstFormat) { case AV_PIX_FMT_NV12: case AV_PIX_FMT_NV24: