From patchwork Tue May 7 16:54:09 2024 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: uk7b@foxmail.com X-Patchwork-Id: 48638 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6a20:9c99:b0:1af:836d:81b3 with SMTP id mj25csp38753pzb; Tue, 7 May 2024 09:55:34 -0700 (PDT) X-Forwarded-Encrypted: i=2; AJvYcCVeqV1QRqITMFbwDi65pI14ROVO2jBvkSPtjI7n8g1t8QVg616lBjcr53nCMr3bXFk9dqKkpx8QaN1tcqWCvnDEE2Q6Jd05ivAYsw== X-Google-Smtp-Source: AGHT+IFfqxK+rvfW91+PjIVDqrbt4Y7VNU62cj2c1Uw3RLNdpGwzg+pIHBJReB+0OSJ68RTjmiXV X-Received: by 2002:a50:aad1:0:b0:572:9dbf:1538 with SMTP id 4fb4d7f45d1cf-5731da81838mr196023a12.31.1715100934169; Tue, 07 May 2024 09:55:34 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1715100934; cv=none; d=google.com; s=arc-20160816; b=vVMPeOpjmklgbByJqTlQim0nLP2/xQ6btFNtL/Pem/gE1vNmHUX3R7fq6Z5ONMVYAW eC8Au5ZnIVn+Map1jGFuyZ+SAgwjwD4KH4MzZNhbWgsY3jrKd7G4YEOnOK22X/AAaqvy wpPLkn8HSLGgDfzpbkfdiN5I6XLb7rNqr4cgacqggYsF1ctMaKsEsX35xJILtqhC6fsl eu/83G6pDfa6rUiglYWiiL9Ysx/uC6pGJSWMpJRHYr8FVuZR1AcUorpiHHhfTkd32EOr vHip+1imxcjR01cwbZBg9wxiGIJNAwd80Jb9sUTfTuOg9q3ih8dcJwGw12Tehsa49mXc T64g== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:mime-version:references:in-reply-to:date :to:from:message-id:dkim-signature:delivered-to; bh=tUonxl1l7MzN0m3+CQY4+JsYO6HPuEmENgSa3Ck9Rxg=; fh=D0bFwGkf4X22/D/bfeDVrXKIx7S6kcXsNzy10j8ORbQ=; b=T8sRziiOdr2e8L0Bga9YS+ArB3+Yh2NDRbuE4WAth2c8j4gDhwLv9HkjcH1op3dK67 Nh4pvLFmNM51kvzZinnLUuUcuZnY3bOxJasd2Dzsv53BdCp61sbFwT/bBxgYldFw7wKe nzTq3heGkdVHPcjc9qS5ZtmT4y3JoMx6L2XwkDiX1hVI3hOXlI2Ij1xoeUJX3j9WMSdk 4+HTQ2yhatjNTmCKfBWmxIVEdqzpeDZR4Ruuv7XwgWGu2qoxmlOUzJDx16jU5Pgbk8Bl ghRu8IiFCstnh/X3lck8yo8ikyv2uh9tmEb4siAWXJSYcgX5JSrl/39ata7ggYazye7B a1kQ==; dara=google.com ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@foxmail.com header.s=s201512 header.b="XE/S2ljt"; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=foxmail.com Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id m17-20020aa7c491000000b005721240c40bsi6063263edq.492.2024.05.07.09.55.31; Tue, 07 May 2024 09:55:34 -0700 (PDT) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@foxmail.com header.s=s201512 header.b="XE/S2ljt"; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=foxmail.com Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 4268768D7CC; Tue, 7 May 2024 19:54:48 +0300 (EEST) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from out162-62-58-216.mail.qq.com (out162-62-58-216.mail.qq.com [162.62.58.216]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 99A5268D752 for ; Tue, 7 May 2024 19:54:38 +0300 (EEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=foxmail.com; s=s201512; t=1715100870; bh=wRvKYaw9HlWJZ4GtcmW7GqvyNO9M2pAUARalWKn/mXA=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=XE/S2ljtRyzahgaNcqfTgwv42G1TYaqOJz3mTB5suiPRaaW8WrZcDm2ZkSbFjTchY luOTxkrKcyddp8iAXVB1tXFMnVbbJ/DMP8zs/IS9zWyRmuVNIYWofNfF8dMOjGbVwm CrC/K5RFMC8Rn5ihzsMTRBvW57MUNujq29xgB0gk= Received: from localhost.localdomain ([42.56.223.122]) by newxmesmtplogicsvrszb9-1.qq.com (NewEsmtp) with SMTP id D99244C2; Wed, 08 May 2024 00:54:25 +0800 X-QQ-mid: xmsmtpt1715100869trrmxyrdg Message-ID: X-QQ-XMAILINFO: NCmjBvJFq6XNJhLqOAHxhs7bT2tF0aTZQqtcy2qwNEJehZotsDc+i4cZxHCsaW 06rVYhUUbv2GPPSLd7Kjf6JerWFKYbxRNNKaW+lcVP1B8+QQTyr4qg1+eWua6d1LSdmY9v97s8fA 5xAbwY+dcfLs89XTmZqF2ZK2aXU0+eZTdMcoPORlPm8tK4LbFag3uNogDfF4DBwbZayaIz2+c+dm oK+A5zBO56yvaa1kdyKFWn0zY0GykyzxtxxCUKEIkKz40RjTKId04o2gB5PNzI86Ev6FPbLmdG05 87gEx7y4OFU3xp1BIPME7BOfOiOoU6dKz6amIjAp76ZVWolrdINLis/vQq/hyQ6nO+JmlbzklgTA omlPJ7U0M1xioLT89clDxXB4uJcnlkH0xwSn6PR/QEXzG5fa3sxAlEQOTaX+lV8WGpseDvn//BDG ndMoX1PsneUy/VJg6hAcJq9mbKezfg1HrUbBH9WB4xdibJRqf942ItA0v887yKcckpAhbLD7+ssH m7I/mQaIBgzpFT9HMrF+IZPX7NqavNn8yQOtoBLsNfXS3VdusGudVPALVkr4kaTJ7k9c1oTcbYM9 BQUcNBKZXRhyMuCMriuD70xemQ9W9fMJk7h52I3BZ6yayHt8B8lyjUZ1quQjQKZ591M/UDLnM8AA vH4wp0bOE5wHuxzaNrsq46gkqPZwk8n6NSxXoyQc+X2eVkU979masVFMdaZMeLlOkuozhFSr/8im 4vb9pgr0g2t5C1B55UG9l4duCbVlQX3H2sGjwRTHr0lQTVsZKC/7NS8e55LMlxJvGf6DhRBJ9Lz/ 4vfR5lBDze8+LWiGIs3F4jm89Pqmbw02zkNezwFn8dCHoNi1usTWLuWkzVVdUgUQpqbRSfD/jhAY udNO9ztPQw/ski/L9nKvisaFTUAi6qrMqVUrY4cMC+JH4y5E9jVaE4pfD9sKnO3Q== X-QQ-XMRINFO: NI4Ajvh11aEj8Xl/2s1/T8w= From: uk7b@foxmail.com To: ffmpeg-devel@ffmpeg.org Date: Wed, 8 May 2024 00:54:09 +0800 X-OQ-MSGID: <20240507165412.1306563-6-uk7b@foxmail.com> X-Mailer: git-send-email 2.45.0 In-Reply-To: <20240507165412.1306563-1-uk7b@foxmail.com> References: <20240507165412.1306563-1-uk7b@foxmail.com> MIME-Version: 1.0 Subject: [FFmpeg-devel] [PATCH v4 6/9] lavc/vp8dsp: R-V V put_epel hv X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: sunyuechi Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: OVydYmYFl+wD From: sunyuechi C908: vp8_put_epel4_h4v4_c: 20.0 vp8_put_epel4_h4v4_rvv_i32: 11.0 vp8_put_epel4_h4v6_c: 25.2 vp8_put_epel4_h4v6_rvv_i32: 13.5 vp8_put_epel4_h6v4_c: 22.2 vp8_put_epel4_h6v4_rvv_i32: 14.5 vp8_put_epel4_h6v6_c: 29.0 vp8_put_epel4_h6v6_rvv_i32: 15.7 vp8_put_epel8_h4v4_c: 73.0 vp8_put_epel8_h4v4_rvv_i32: 22.2 vp8_put_epel8_h4v6_c: 90.5 vp8_put_epel8_h4v6_rvv_i32: 26.7 vp8_put_epel8_h6v4_c: 85.0 vp8_put_epel8_h6v4_rvv_i32: 27.2 vp8_put_epel8_h6v6_c: 104.7 vp8_put_epel8_h6v6_rvv_i32: 29.5 vp8_put_epel16_h4v4_c: 145.5 vp8_put_epel16_h4v4_rvv_i32: 26.5 vp8_put_epel16_h4v6_c: 190.7 vp8_put_epel16_h4v6_rvv_i32: 47.5 vp8_put_epel16_h6v4_c: 173.7 vp8_put_epel16_h6v4_rvv_i32: 33.2 vp8_put_epel16_h6v6_c: 222.2 vp8_put_epel16_h6v6_rvv_i32: 35.5 --- libavcodec/riscv/vp8dsp_init.c | 13 ++++ libavcodec/riscv/vp8dsp_rvv.S | 123 +++++++++++++++++++++++++++------ 2 files changed, 115 insertions(+), 21 deletions(-) diff --git a/libavcodec/riscv/vp8dsp_init.c b/libavcodec/riscv/vp8dsp_init.c index dc3e087f01..463c8fa0a2 100644 --- a/libavcodec/riscv/vp8dsp_init.c +++ b/libavcodec/riscv/vp8dsp_init.c @@ -97,6 +97,19 @@ av_cold void ff_vp78dsp_init_riscv(VP8DSPContext *c) c->put_vp8_epel_pixels_tab[0][1][0] = ff_put_vp8_epel16_v4_rvv; c->put_vp8_epel_pixels_tab[1][1][0] = ff_put_vp8_epel8_v4_rvv; c->put_vp8_epel_pixels_tab[2][1][0] = ff_put_vp8_epel4_v4_rvv; + + c->put_vp8_epel_pixels_tab[0][2][2] = ff_put_vp8_epel16_h6v6_rvv; + c->put_vp8_epel_pixels_tab[1][2][2] = ff_put_vp8_epel8_h6v6_rvv; + c->put_vp8_epel_pixels_tab[2][2][2] = ff_put_vp8_epel4_h6v6_rvv; + c->put_vp8_epel_pixels_tab[0][2][1] = ff_put_vp8_epel16_h4v6_rvv; + c->put_vp8_epel_pixels_tab[1][2][1] = ff_put_vp8_epel8_h4v6_rvv; + c->put_vp8_epel_pixels_tab[2][2][1] = ff_put_vp8_epel4_h4v6_rvv; + c->put_vp8_epel_pixels_tab[0][1][1] = ff_put_vp8_epel16_h4v4_rvv; + c->put_vp8_epel_pixels_tab[1][1][1] = ff_put_vp8_epel8_h4v4_rvv; + c->put_vp8_epel_pixels_tab[2][1][1] = ff_put_vp8_epel4_h4v4_rvv; + c->put_vp8_epel_pixels_tab[0][1][2] = ff_put_vp8_epel16_h6v4_rvv; + c->put_vp8_epel_pixels_tab[1][1][2] = ff_put_vp8_epel8_h6v4_rvv; + c->put_vp8_epel_pixels_tab[2][1][2] = ff_put_vp8_epel4_h6v4_rvv; } #endif #endif diff --git a/libavcodec/riscv/vp8dsp_rvv.S b/libavcodec/riscv/vp8dsp_rvv.S index 4d7a9f6a2d..fba72f8c15 100644 --- a/libavcodec/riscv/vp8dsp_rvv.S +++ b/libavcodec/riscv/vp8dsp_rvv.S @@ -161,26 +161,26 @@ const subpel_filters .byte 0, -1, 12, 123, -6, 0 endconst -.macro epel_filter size type - lla t2, subpel_filters +.macro epel_filter size type regtype + lla \regtype\()2, subpel_filters .ifc \type,v - addi t0, a6, -1 + addi \regtype\()0, a6, -1 .else - addi t0, a5, -1 + addi \regtype\()0, a5, -1 .endif - li t1, 6 - mul t0, t0, t1 - add t0, t0, t2 + li \regtype\()1, 6 + mul \regtype\()0, \regtype\()0, \regtype\()1 + add \regtype\()0, \regtype\()0, \regtype\()2 .irp n 1,2,3,4 - lb t\n, \n(t0) + lb \regtype\n, \n(\regtype\()0) .endr .ifc \size,6 - lb t5, 5(t0) - lb t0, (t0) + lb \regtype\()5, 5(\regtype\()0) + lb \regtype\()0, (\regtype\()0) .endif .endm -.macro epel_load dst len size type +.macro epel_load dst len size type from_mem regtype .ifc \type,v mv a5, a3 .else @@ -189,24 +189,35 @@ endconst sub t6, a2, a5 add a7, a2, a5 +.if \from_mem vle8.v v24, (a2) vle8.v v22, (t6) vle8.v v26, (a7) add a7, a7, a5 vle8.v v28, (a7) - vwmulu.vx v16, v24, t2 - vwmulu.vx v20, v26, t3 + vwmulu.vx v16, v24, \regtype\()2 + vwmulu.vx v20, v26, \regtype\()3 .ifc \size,6 sub t6, t6, a5 add a7, a7, a5 vle8.v v24, (t6) vle8.v v26, (a7) - vwmaccu.vx v16, t0, v24 - vwmaccu.vx v16, t5, v26 + vwmaccu.vx v16, \regtype\()0, v24 + vwmaccu.vx v16, \regtype\()5, v26 +.endif + vwmaccsu.vx v16, \regtype\()1, v22 + vwmaccsu.vx v16, \regtype\()4, v28 +.else + vwmulu.vx v16, v4, \regtype\()2 + vwmulu.vx v20, v6, \regtype\()3 + .ifc \size,6 + vwmaccu.vx v16, \regtype\()0, v0 + vwmaccu.vx v16, \regtype\()5, v10 + .endif + vwmaccsu.vx v16, \regtype\()1, v2 + vwmaccsu.vx v16, \regtype\()4, v8 .endif li t6, 64 - vwmaccsu.vx v16, t1, v22 - vwmaccsu.vx v16, t4, v28 vwadd.wx v16, v16, t6 vsetvlstatic16 \len vwadd.vv v24, v16, v20 @@ -216,18 +227,18 @@ endconst vnclipu.wi \dst, v24, 0 .endm -.macro epel_load_inc dst len size type - epel_load \dst \len \size \type +.macro epel_load_inc dst len size type from_mem regtype + epel_load \dst \len \size \type \from_mem \regtype add a2, a2, a3 .endm .macro epel len size type func ff_put_vp8_epel\len\()_\type\()\size\()_rvv, zve32x - epel_filter \size \type + epel_filter \size \type t vsetvlstatic8 \len 1: addi a4, a4, -1 - epel_load_inc v30 \len \size \type + epel_load_inc v30 \len \size \type 1 t vse8.v v30, (a0) add a0, a0, a1 bnez a4, 1b @@ -236,6 +247,72 @@ func ff_put_vp8_epel\len\()_\type\()\size\()_rvv, zve32x endfunc .endm +.macro epel_hv len hsize vsize +func ff_put_vp8_epel\len\()_h\hsize\()v\vsize\()_rvv, zve32x +#if __riscv_xlen == 64 + addi sp, sp, -48 + .irp n 0,1,2,3,4,5 + sd s\n, \n\()<<3(sp) + .endr +#else + addi sp, sp, -24 + .irp n 0,1,2,3,4,5 + sw s\n, \n\()<<2(sp) + .endr +#endif + sub a2, a2, a3 + epel_filter \hsize h t + epel_filter \vsize v s + vsetvlstatic8 \len +.if \hsize == 6 || \vsize == 6 + sub a2, a2, a3 + epel_load_inc v0 \len \hsize h 1 t +.endif + epel_load_inc v2 \len \hsize h 1 t + epel_load_inc v4 \len \hsize h 1 t + epel_load_inc v6 \len \hsize h 1 t + epel_load_inc v8 \len \hsize h 1 t +.if \hsize == 6 || \vsize == 6 + epel_load_inc v10 \len \hsize h 1 t +.endif + addi a4, a4, -1 +1: + addi a4, a4, -1 + epel_load v30 \len \vsize v 0 s + vse8.v v30, (a0) +.if \hsize == 6 || \vsize == 6 + vmv.v.v v0, v2 +.endif + vmv.v.v v2, v4 + vmv.v.v v4, v6 + vmv.v.v v6, v8 +.if \hsize == 6 || \vsize == 6 + vmv.v.v v8, v10 + epel_load_inc v10 \len \hsize h 1 t +.else + epel_load_inc v8 \len 4 h 1 t +.endif + add a0, a0, a1 + bnez a4, 1b + epel_load v30 \len \vsize v 0 s + vse8.v v30, (a0) + +#if __riscv_xlen == 64 + .irp n 0,1,2,3,4,5 + ld s\n, \n\()<<3(sp) + .endr + addi sp, sp, 48 +#else + .irp n 0,1,2,3,4,5 + lw s\n, \n\()<<2(sp) + .endr + addi sp, sp, 24 +#endif + + ret +endfunc +.endm + .irp len 16,8,4 put_vp8_bilin_h_v \len h a5 put_vp8_bilin_h_v \len v a6 @@ -244,4 +321,8 @@ epel \len 6 h epel \len 4 h epel \len 6 v epel \len 4 v +epel_hv \len 6 6 +epel_hv \len 4 4 +epel_hv \len 6 4 +epel_hv \len 4 6 .endr