From patchwork Fri Oct 22 08:52:28 2021 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Wu Jianhua X-Patchwork-Id: 31210 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6602:2084:0:0:0:0 with SMTP id a4csp1557108ioa; Fri, 22 Oct 2021 01:53:19 -0700 (PDT) X-Google-Smtp-Source: ABdhPJwrzaId0f0cMjzpf2V46k5FWPzyYd7DWY2r91miF3sMz7NWUnJFt4EQ+B54sf+XbU+On+Lq X-Received: by 2002:aa7:cd88:: with SMTP id x8mr14728497edv.203.1634892799260; Fri, 22 Oct 2021 01:53:19 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1634892799; cv=none; d=google.com; s=arc-20160816; b=GxwqMGb20yo508YoYpnDPTDMwGpKeUBDJtVFBn1oiDFpkw+bbsvhl7rjnWrZ1z2JwB bKOPITuWwPZYwxdc+7dDBcNQiShf3yl/xwwM2GE6UpLAYACfSuC+Om0UNPLU2SwfyXsC xbu8HVgj1KPKcZPYXUfAV2AzIyifC8xXWZwGZF/7pyVK9IWOUAPI7RH87MFn8ITxx5gq TK8aVUypQ18i9W8Y8qOuIkUL6+LtkK7OKoiSQ7At/90sN03AA5GgcIVulfCI2dWckrqK CnwvK7n2tPc0snmbEN6YEBj3fvglObkvg6N4T8+GVCuBuY9fjdZFSnxdW7KgWI/F7tRq ohLQ== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:mime-version:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:message-id:date:to:from:delivered-to; bh=THrkBvsyZIPzavz0aKVbbjYepkw/1GuokGVqjJkrRNk=; b=kJCm7S1YwEJ64fCKY0nZTKqknYo6McVmSEbFTp8AxO7AAIewCj7wiQwH5GQbAjRxSs kVAsyXmngcNOYEFwOw5Ai6M+naJmL+DjZ5Jxf7EjJd7gGHCwXPYgzVs8YjTVKuHAU2ht qwInbsj351h76tKJOVx1RZw2DWIIrwLmTjGlzC0nMl8+7FA42KpfBJF4K/kYadRSolY2 quHIniz9Yyi3T7mrvrf93PK86W8pCevRtqwIzTPUUvixULPYmotXytLSWl0saAOc6Tvk CdwF4R7uN2us326pKlw1Sx6RkWLki57HgJ+Xvx6bX++AMm0pflsTn48AZYHSZcR+1MYP WL2A== ARC-Authentication-Results: i=1; mx.google.com; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=intel.com Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id hp24si15455925ejc.400.2021.10.22.01.53.18; Fri, 22 Oct 2021 01:53:19 -0700 (PDT) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=intel.com Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 6342068A71A; Fri, 22 Oct 2021 11:53:08 +0300 (EEST) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mga18.intel.com (mga18.intel.com [134.134.136.126]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 6AAF768A242 for ; Fri, 22 Oct 2021 11:53:01 +0300 (EEST) X-IronPort-AV: E=McAfee;i="6200,9189,10144"; a="216176982" X-IronPort-AV: E=Sophos;i="5.87,172,1631602800"; d="scan'208";a="216176982" Received: from orsmga008.jf.intel.com ([10.7.209.65]) by orsmga106.jf.intel.com with ESMTP/TLS/ECDHE-RSA-AES256-GCM-SHA384; 22 Oct 2021 01:52:51 -0700 X-ExtLoop1: 1 X-IronPort-AV: E=Sophos;i="5.87,172,1631602800"; d="scan'208";a="495586607" Received: from otc-skl-e5-server.sh.intel.com ([10.239.43.106]) by orsmga008.jf.intel.com with ESMTP; 22 Oct 2021 01:52:49 -0700 From: Wu Jianhua To: ffmpeg-devel@ffmpeg.org Date: Fri, 22 Oct 2021 16:52:28 +0800 Message-Id: <20211022085231.93931-1-jianhua.wu@intel.com> X-Mailer: git-send-email 2.17.1 Subject: [FFmpeg-devel] [PATCH 1/4] avfilter/x86/vf_exposure: add x86 SIMD optimization X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Wu Jianhua MIME-Version: 1.0 Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: FC9JHz/gWxQQ Performance data(Less is better): exposure_c: 857394 exposure_asm: 327589 Signed-off-by: Wu Jianhua --- libavfilter/exposure.h | 36 +++++++++++++++++++ libavfilter/vf_exposure.c | 36 +++++++++---------- libavfilter/x86/Makefile | 2 ++ libavfilter/x86/vf_exposure.asm | 55 ++++++++++++++++++++++++++++++ libavfilter/x86/vf_exposure_init.c | 36 +++++++++++++++++++ 5 files changed, 147 insertions(+), 18 deletions(-) create mode 100644 libavfilter/exposure.h create mode 100644 libavfilter/x86/vf_exposure.asm create mode 100644 libavfilter/x86/vf_exposure_init.c diff --git a/libavfilter/exposure.h b/libavfilter/exposure.h new file mode 100644 index 0000000000..e76a517826 --- /dev/null +++ b/libavfilter/exposure.h @@ -0,0 +1,36 @@ +/* + * This file is part of FFmpeg. + * + * FFmpeg is free software; you can redistribute it and/or + * modify it under the terms of the GNU Lesser General Public + * License as published by the Free Software Foundation; either + * version 2.1 of the License, or (at your option) any later version. + * + * FFmpeg is distributed in the hope that it will be useful, + * but WITHOUT ANY WARRANTY; without even the implied warranty of + * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU + * Lesser General Public License for more details. + * + * You should have received a copy of the GNU Lesser General Public + * License along with FFmpeg; if not, write to the Free Software + * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA + */ + +#ifndef AVFILTER_EXPOSURE_H +#define AVFILTER_EXPOSURE_H +#include "avfilter.h" + +typedef struct ExposureContext { + const AVClass *class; + + float exposure; + float black; + float scale; + + void (*exposure_func)(float *ptr, int length, float black, float scale); +} ExposureContext; + +void ff_exposure_init(ExposureContext *s); +void ff_exposure_init_x86(ExposureContext *s); + +#endif diff --git a/libavfilter/vf_exposure.c b/libavfilter/vf_exposure.c index 108fba7930..045ae710d3 100644 --- a/libavfilter/vf_exposure.c +++ b/libavfilter/vf_exposure.c @@ -26,23 +26,20 @@ #include "formats.h" #include "internal.h" #include "video.h" +#include "exposure.h" -typedef struct ExposureContext { - const AVClass *class; - - float exposure; - float black; +static void exposure_c(float *ptr, int length, float black, float scale) +{ + int i; - float scale; - int (*do_slice)(AVFilterContext *s, void *arg, - int jobnr, int nb_jobs); -} ExposureContext; + for (i = 0; i < length; i++) + ptr[i] = (ptr[i] - black) * scale; +} static int exposure_slice(AVFilterContext *ctx, void *arg, int jobnr, int nb_jobs) { ExposureContext *s = ctx->priv; AVFrame *frame = arg; - const int width = frame->width; const int height = frame->height; const int slice_start = (height * jobnr) / nb_jobs; const int slice_end = (height * (jobnr + 1)) / nb_jobs; @@ -52,24 +49,27 @@ static int exposure_slice(AVFilterContext *ctx, void *arg, int jobnr, int nb_job for (int p = 0; p < 3; p++) { const int linesize = frame->linesize[p] / 4; float *ptr = (float *)frame->data[p] + slice_start * linesize; - for (int y = slice_start; y < slice_end; y++) { - for (int x = 0; x < width; x++) - ptr[x] = (ptr[x] - black) * scale; - - ptr += linesize; - } + s->exposure_func(ptr, linesize * (slice_end - slice_start), black, scale); } return 0; } +void ff_exposure_init(ExposureContext *s) +{ + s->exposure_func = exposure_c; + + if (ARCH_X86) + ff_exposure_init_x86(s); +} + static int filter_frame(AVFilterLink *inlink, AVFrame *frame) { AVFilterContext *ctx = inlink->dst; ExposureContext *s = ctx->priv; s->scale = 1.f / (exp2f(-s->exposure) - s->black); - ff_filter_execute(ctx, s->do_slice, frame, NULL, + ff_filter_execute(ctx, exposure_slice, frame, NULL, FFMIN(frame->height, ff_filter_get_nb_threads(ctx))); return ff_filter_frame(ctx->outputs[0], frame); @@ -80,7 +80,7 @@ static av_cold int config_input(AVFilterLink *inlink) AVFilterContext *ctx = inlink->dst; ExposureContext *s = ctx->priv; - s->do_slice = exposure_slice; + ff_exposure_init(s); return 0; } diff --git a/libavfilter/x86/Makefile b/libavfilter/x86/Makefile index a29941eaeb..e84a388aa5 100644 --- a/libavfilter/x86/Makefile +++ b/libavfilter/x86/Makefile @@ -8,6 +8,7 @@ OBJS-$(CONFIG_BWDIF_FILTER) += x86/vf_bwdif_init.o OBJS-$(CONFIG_COLORSPACE_FILTER) += x86/colorspacedsp_init.o OBJS-$(CONFIG_CONVOLUTION_FILTER) += x86/vf_convolution_init.o OBJS-$(CONFIG_EQ_FILTER) += x86/vf_eq_init.o +OBJS-$(CONFIG_EXPOSURE_FILTER) += x86/vf_exposure_init.o OBJS-$(CONFIG_FSPP_FILTER) += x86/vf_fspp_init.o OBJS-$(CONFIG_GBLUR_FILTER) += x86/vf_gblur_init.o OBJS-$(CONFIG_GRADFUN_FILTER) += x86/vf_gradfun_init.o @@ -49,6 +50,7 @@ X86ASM-OBJS-$(CONFIG_BWDIF_FILTER) += x86/vf_bwdif.o X86ASM-OBJS-$(CONFIG_COLORSPACE_FILTER) += x86/colorspacedsp.o X86ASM-OBJS-$(CONFIG_CONVOLUTION_FILTER) += x86/vf_convolution.o X86ASM-OBJS-$(CONFIG_EQ_FILTER) += x86/vf_eq.o +X86ASM-OBJS-$(CONFIG_EXPOSURE_FILTER) += x86/vf_exposure.o X86ASM-OBJS-$(CONFIG_FRAMERATE_FILTER) += x86/vf_framerate.o X86ASM-OBJS-$(CONFIG_FSPP_FILTER) += x86/vf_fspp.o X86ASM-OBJS-$(CONFIG_GBLUR_FILTER) += x86/vf_gblur.o diff --git a/libavfilter/x86/vf_exposure.asm b/libavfilter/x86/vf_exposure.asm new file mode 100644 index 0000000000..3351c6fb3b --- /dev/null +++ b/libavfilter/x86/vf_exposure.asm @@ -0,0 +1,55 @@ +;***************************************************************************** +;* x86-optimized functions for exposure filter +;* +;* This file is part of FFmpeg. +;* +;* FFmpeg is free software; you can redistribute it and/or +;* modify it under the terms of the GNU Lesser General Public +;* License as published by the Free Software Foundation; either +;* version 2.1 of the License, or (at your option) any later version. +;* +;* FFmpeg is distributed in the hope that it will be useful, +;* but WITHOUT ANY WARRANTY; without even the implied warranty of +;* MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU +;* Lesser General Public License for more details. +;* +;* You should have received a copy of the GNU Lesser General Public +;* License along with FFmpeg; if not, write to the Free Software +;* Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA +;****************************************************************************** + +%include "libavutil/x86/x86util.asm" + +SECTION .text + +;******************************************************************************* +; void ff_exposure(float *ptr, int length, float black, float scale); +;******************************************************************************* +%macro EXPOSURE 0 +cglobal exposure, 2, 2, 4, ptr, length, black, scale + movsxdifnidn lengthq, lengthd +%if WIN64 + VBROADCASTSS m0, xmm2 + VBROADCASTSS m1, xmm3 +%else + VBROADCASTSS m0, xmm0 + VBROADCASTSS m1, xmm1 +%endif + +.loop: + movu m2, [ptrq] + subps m2, m2, m0 + mulps m2, m2, m1 + movu [ptrq], m2 + add ptrq, mmsize + sub lengthq, mmsize/4 + + jg .loop + + RET +%endmacro + +%if ARCH_X86_64 +INIT_XMM sse +EXPOSURE +%endif diff --git a/libavfilter/x86/vf_exposure_init.c b/libavfilter/x86/vf_exposure_init.c new file mode 100644 index 0000000000..de1b360f6c --- /dev/null +++ b/libavfilter/x86/vf_exposure_init.c @@ -0,0 +1,36 @@ +/* + * This file is part of FFmpeg. + * + * FFmpeg is free software; you can redistribute it and/or + * modify it under the terms of the GNU Lesser General Public + * License as published by the Free Software Foundation; either + * version 2.1 of the License, or (at your option) any later version. + * + * FFmpeg is distributed in the hope that it will be useful, + * but WITHOUT ANY WARRANTY; without even the implied warranty of + * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU + * Lesser General Public License for more details. + * + * You should have received a copy of the GNU Lesser General Public + * License along with FFmpeg; if not, write to the Free Software + * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA + */ + +#include "config.h" + +#include "libavutil/attributes.h" +#include "libavutil/cpu.h" +#include "libavutil/x86/cpu.h" +#include "libavfilter/exposure.h" + +void ff_exposure_sse(float *ptr, int length, float black, float scale); + +av_cold void ff_exposure_init_x86(ExposureContext *s) +{ + int cpu_flags = av_get_cpu_flags(); + +#if ARCH_X86_64 + if (EXTERNAL_SSE(cpu_flags)) + s->exposure_func = ff_exposure_sse; +#endif +}