From patchwork Sat Jun 22 15:58:06 2024 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: uk7b@foxmail.com X-Patchwork-Id: 50084 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a59:ae71:0:b0:482:c625:d099 with SMTP id w17csp1137731vqz; Sat, 22 Jun 2024 08:58:36 -0700 (PDT) X-Forwarded-Encrypted: i=2; AJvYcCXIyz+C4DVqco6WUKgpU9oUYXvxFGrDO73fbmxypXtyUf+xI4tlNQju5Z55fqkAkRSPR/WZCyjMzom9dva/hc+i1m/66RFCYKHL/Q== X-Google-Smtp-Source: AGHT+IHNbPv13hsglZ4iH1CJm+50i2V8TBxauPAcZ/LDmgLDX96CQezUIbPx/d39L5Q4vnVsJ7CE X-Received: by 2002:a17:906:5957:b0:a6f:dffc:54b6 with SMTP id a640c23a62f3a-a7242e142d8mr39550966b.73.1719071915937; Sat, 22 Jun 2024 08:58:35 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1719071915; cv=none; d=google.com; s=arc-20160816; b=Urc1gXtlLfwTVto56TJkmoacYY5M437ophp+UcfwF1yj9U8fP1tJhG8rrI58zL7XNc raB9ONqYG/bCt9YLPEijYJBAPUQhk8mj8F2WyZRT1bDuCQDDRGgvgdfT/O7EKFEcFslJ X4ZitBtY4t86w5PMQxg2QPnthWsAmp/wR1g4f0LCnxI1XnUnbQDAbOstkzsNKDwdG2LX +ONnquI0po2jruLAg7IRGRK6hQqNklX1+iFUYI2rfCuKmVR8smPqJyK12gEHirChmpdS byDYrGiWXo3chqiNnpaVLok0ADR2vRQaaJY77I5IsPAr7BVIvJZU4xKhOe1XNlMD+S4A uiqA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:mime-version:references:in-reply-to:date :to:from:message-id:dkim-signature:delivered-to; bh=YklpZcwOUkIRa5kAZWZ4bawquj8CeHwcsOKRnmv3ei0=; fh=D0bFwGkf4X22/D/bfeDVrXKIx7S6kcXsNzy10j8ORbQ=; b=PcaQkAmUwlM4lq0AZwtbdsK0RXPZgSZRx/vyHcXBqDoNGFgJlCrGaPqsztcvXOcbka 91CnsU0SEGsKn4FRAEI0iQZEIGW4lI2fwnHGt1yoKEfg4Ekk6wXQ1ozYDkmNcBEw+GhC a9opsQjcpra0C7++GoJJR8bNxUqTPyuHFdqP3GN7F0ROqN1ONc2n0kfMGPfvpBMBTNAz FHyk89QkCT/2T3TkPuHPrqqWDSb5yJ79eLevQFIXBsmPSpiBeIM3iieCBcylC/C91osw YJDqdBr3c8RZRpayo2cNfniV8C+SGNPplkHvgq2FeINhx3uhVHCeUGGnlBJgwvp+29MD 7AOQ==; dara=google.com ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@foxmail.com header.s=s201512 header.b=LsjpPgG3; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=foxmail.com Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id a640c23a62f3a-a6fcf5953f7si196190766b.890.2024.06.22.08.58.35; Sat, 22 Jun 2024 08:58:35 -0700 (PDT) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@foxmail.com header.s=s201512 header.b=LsjpPgG3; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=foxmail.com Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 5A96B68D406; Sat, 22 Jun 2024 18:58:31 +0300 (EEST) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from out162-62-57-252.mail.qq.com (out162-62-57-252.mail.qq.com [162.62.57.252]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id CA74568D128 for ; Sat, 22 Jun 2024 18:58:22 +0300 (EEST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=foxmail.com; s=s201512; t=1719071894; bh=AVJeVJHkXdfbFIdqYlxmSDctI++1FXd0dwd70TifB3U=; h=From:To:Cc:Subject:Date:In-Reply-To:References; b=LsjpPgG3FNvD607cOHAKoNldOaJysPdh3piLnp7i/M+cOr5ohPyL6Vmdjrikanuw+ pQK+s3844CNVnmgi2c8MbbwA0CIhThHwRxdVKNk4Yye2ccdk7YkOVUnnU48PSY2NQe +EBDoUo2x169KovwVusgTRnU5+IStJFc1zVLAdo4= Received: from localhost.localdomain ([116.139.97.96]) by newxmesmtplogicsvrsza10-0.qq.com (NewEsmtp) with SMTP id E8A868C7; Sat, 22 Jun 2024 23:58:10 +0800 X-QQ-mid: xmsmtpt1719071893trg7hmy9c Message-ID: X-QQ-XMAILINFO: NhpLzBn2I3Xwy8hCfkUzzj5jKtoinMIxIT90bbZYZLsZ8Cry9io+9XI4QXQxJR b42GbLOSOX0viIPdRrJZ1KXi+TYFxoQiA2oBiPepl6iGmFX3Rwt5hVRJtUzNEQpJdFZ3bIMHMdAK n+S0HcIDRO8/N88UjN5bPmvPyGFeji0h2f8qR42mqaKFIElWEkeKuzB7xOje/4ja3iRNO5vyHx2q tCvqPdks1NvO2leiCI76kkFLv1ncppUgiZQUv+bhzEDy46rV/8ujL8vF5AI3Rk3Et9LVr2pjJrST tcGgObFeDfafj8zvIkiRMYML4OaCAipr89xU4Feli3o24zJFfa3ychG5u6XUfo05m6GSmcP8uAeY 1UUU6wxDbT1OVRPh5oBkbVdv9dHuS5pfLDR4EMcLQ4IDXSrSxJIT3ADxEpLFyUZczCykQovh7Xxe TISRZw2+H3K1zQ0IfwhUX00SGlv2asbKPr19j8ftk4uQ8r2y+CFBrHztw7xTjgWxLtTLxJOce/3N e5JmxeMYvfheO9alzX4zxnJ6Csg+2lVMUAKkI/8aNW4GO37wSLm4u+ut3/LU7G2edwg5K4QFWN9E ngFgbQXdpaZSlzxE0U8VqNEoV8OnV933WoWGFv98h8dKOdM6S65G7jtuF5wdH+K1vm4E35yp9o8+ 9rN6Xbxq7//EXXRob2Ir9C1bnyscojHUM24WOG1Spx9BMLV1ntyX1JceKC+87Vk6XQVlQ9V82xdX WX2R+uuSVhhmcldpwg430RXqL23OSrW8awYcX3uZM3GG0ewpCH5nCgf4V996a3WjxCGZJIhfb/N3 ih8DrZmErRzcYS+9eIPpQcYQkwq0L/YzIbG06A8KMB25x799YkGmqlgI1QH0AMcH+MxXHlksJXA/ LzdEZDOKE6r2eKFo4tCI6SuLhSqvIUmCsy1f716YdX X-QQ-XMRINFO: NI4Ajvh11aEj8Xl/2s1/T8w= From: uk7b@foxmail.com To: ffmpeg-devel@ffmpeg.org Date: Sat, 22 Jun 2024 23:58:06 +0800 X-OQ-MSGID: <20240622155806.3191984-4-uk7b@foxmail.com> X-Mailer: git-send-email 2.45.2 In-Reply-To: <20240622155806.3191984-1-uk7b@foxmail.com> References: <20240622155806.3191984-1-uk7b@foxmail.com> MIME-Version: 1.0 Subject: [FFmpeg-devel] [PATCH 4/4] lavc/vp8dsp: R-V V loop_filter X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: sunyuechi Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: 0wnF9nW3ZfpR From: sunyuechi C908 X60 vp8_loop_filter8uv_v_c : 13.7 11.7 vp8_loop_filter8uv_v_rvv_i32 : 7.7 6.2 vp8_loop_filter16y_h_c : 12.2 11.2 vp8_loop_filter16y_h_rvv_i32 : 9.5 7.2 vp8_loop_filter16y_v_c : 13.2 12.0 vp8_loop_filter16y_v_rvv_i32 : 5.5 3.7 --- libavcodec/riscv/vp8dsp_init.c | 5 ++- libavcodec/riscv/vp8dsp_rvv.S | 57 ++++++++++++++++++++++++++++++++++ 2 files changed, 61 insertions(+), 1 deletion(-) diff --git a/libavcodec/riscv/vp8dsp_init.c b/libavcodec/riscv/vp8dsp_init.c index 94f78cd84b..72191f558b 100644 --- a/libavcodec/riscv/vp8dsp_init.c +++ b/libavcodec/riscv/vp8dsp_init.c @@ -157,7 +157,10 @@ av_cold void ff_vp8dsp_init_riscv(VP8DSPContext *c) c->vp8_h_loop_filter_simple = ff_vp8_h_loop_filter16_simple_rvv##vlen; \ c->vp8_v_loop_filter16y_inner = ff_vp8_v_loop_filter16_inner_rvv##vlen; \ c->vp8_h_loop_filter16y_inner = ff_vp8_h_loop_filter16_inner_rvv##vlen; \ - c->vp8_v_loop_filter8uv_inner = ff_vp8_v_loop_filter8uv_inner_rvv##vlen; + c->vp8_v_loop_filter8uv_inner = ff_vp8_v_loop_filter8uv_inner_rvv##vlen; \ + c->vp8_v_loop_filter16y = ff_vp8_v_loop_filter16_rvv##vlen; \ + c->vp8_h_loop_filter16y = ff_vp8_h_loop_filter16_rvv##vlen; \ + c->vp8_v_loop_filter8uv = ff_vp8_v_loop_filter8uv_rvv##vlen; int flags = av_get_cpu_flags(); diff --git a/libavcodec/riscv/vp8dsp_rvv.S b/libavcodec/riscv/vp8dsp_rvv.S index ed789ec4fd..98ff389b9a 100644 --- a/libavcodec/riscv/vp8dsp_rvv.S +++ b/libavcodec/riscv/vp8dsp_rvv.S @@ -410,6 +410,33 @@ endfunc vsra.vi v24, v24, 1 // (f1 + 1) >> 1; vadd.vv v8, v18, v24 vsub.vv v10, v20, v24 + .else + li t5, 27 + li t3, 9 + li a7, 18 + vwmul.vx v2, v11, t5 + vwmul.vx v6, v11, t3 + vwmul.vx v4, v11, a7 + vsetvlstatic16 \len, \vlen + li a7, 63 + vzext.vf2 v14, v15 // p2 + vzext.vf2 v24, v10 // q2 + vadd.vx v2, v2, a7 + vadd.vx v4, v4, a7 + vadd.vx v6, v6, a7 + vsra.vi v2, v2, 7 // a0 + vsra.vi v12, v4, 7 // a1 + vsra.vi v6, v6, 7 // a2 + vadd.vv v14, v14, v6 // p2 + a2 + vsub.vv v22, v24, v6 // q2 - a2 + vsub.vv v10, v20, v12 // q1 - a1 + vadd.vv v4, v8, v2 // p0 + a0 + vsub.vv v6, v16, v2 // q0 - a0 + vadd.vv v8, v12, v18 // a1 + p1 + vmax.vx v4, v4, zero + vmax.vx v6, v6, zero + vmax.vx v14, v14, zero + vmax.vx v16, v22, zero .endif vmax.vx v8, v8, zero @@ -430,6 +457,17 @@ endfunc vsse8.v v6, (a6), \stride, v0.t vsse8.v v7, (t4), \stride, v0.t .endif + .if !\inner + vnclipu.wi v14, v14, 0 + vnclipu.wi v16, v16, 0 + .ifc \type,v + vse8.v v14, (t0), v0.t + vse8.v v16, (t6), v0.t + .else + vsse8.v v14, (t0), \stride, v0.t + vsse8.v v16, (t6), \stride, v0.t + .endif + .endif .endif .endm @@ -464,6 +502,25 @@ func ff_vp8_v_loop_filter8uv_inner_rvv\vlen, zve32x filter 8, \vlen, v, 1, 1, a1, a2, a3, a4, a5 ret endfunc + +func ff_vp8_v_loop_filter16_rvv\vlen, zve32x + vsetvlstatic8 16, \vlen + filter 16, \vlen, v, 1, 0, a0, a1, a2, a3, a4 + ret +endfunc + +func ff_vp8_h_loop_filter16_rvv\vlen, zve32x + vsetvlstatic8 16, \vlen + filter 16, \vlen, h, 1, 0, a0, a1, a2, a3, a4 + ret +endfunc + +func ff_vp8_v_loop_filter8uv_rvv\vlen, zve32x + vsetvlstatic8 8, \vlen + filter 8, \vlen, v, 1, 0, a0, a2, a3, a4, a5 + filter 8, \vlen, v, 1, 0, a1, a2, a3, a4, a5 + ret +endfunc .endr .macro bilin_load_h dst mn