From patchwork Mon Dec 20 14:43:12 2021 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 32756 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a6b:cd86:0:0:0:0:0 with SMTP id d128csp4427123iog; Mon, 20 Dec 2021 06:43:28 -0800 (PST) X-Google-Smtp-Source: ABdhPJwJDBEP0EA36ZZ11qeViOGu4PM2DRI/ZTO0Ab4fWpojmRQSjOhBITwfVR5Gsn4+tDW8KcQ4 X-Received: by 2002:a05:6402:158b:: with SMTP id c11mr16375480edv.293.1640011407879; Mon, 20 Dec 2021 06:43:27 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1640011407; cv=none; d=google.com; s=arc-20160816; b=IRT1SBqvssOqnQtKjderjOb7UeRsawLU7B5mU7wRSH3JRd2H0VGlBC3jIrNf0Txj4r 9avwK803rnTwFAMvOZYiAbcxZ8FxkX6cwr5Wivf4euw1YzVU4f3ZQvVDDexS4q9lJmo6 oYIjCoV5OiqRPRjSVuy6vW2TfqxcnkZk54TmNoeEw4yMapSVcEVneBLrf/Js772ektrV ecQQk0vr5HDRmwYxVMlWZEqgbJfCSo5J6tyEHbY/Sy3jve+FBJDRXB20byj1Veo5EDI/ FmJ7F/GBrokWxlh2Hp8oAUmPgmUDGCOhwtSkISENvt6WQoWbrw5L8gyOvPDpEXBGTcMH ZGkA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:references:mime-version :message-id:in-reply-to:date:dkim-signature:delivered-to; bh=bviVPpNOmMTUvm9Wk6Fpj2tpvi4hx+18RdulRdQqGG8=; b=LMS9Dv5R34nBwcL0I6llofVWtyQy5WMJivm49Vke5hgOOUt4iO67DLMLi2UKMxHgEw yS66I60wUTx0cYSrCaq14R4mjPmLmcRGmsn21jZfMyPRJmOMQ/nFEMhkK1boF7COGF3H qtImKMlc2CccdA3PccUDTAg7h74hLS0s2ZjkQl6LwvkQtl3HURNmbnMbb2H2W9DT9aSQ I04Ac9vB1ybfNQ19MM7RcuSY/52WuCaLWg7gWYIY5LrXwY1d0+TN6SJN7Z5+pBoRZl2r mQb1/NAlqG4dD1/UN+jiokN9a1OTx2n87xYO70gujfHJQ1B1AZEzyb2XkC87/DwlR5zm g9Gg== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=SClSBDsa; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id j17si7413121edf.262.2021.12.20.06.43.26; Mon, 20 Dec 2021 06:43:27 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=SClSBDsa; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 7D11668AE86; Mon, 20 Dec 2021 16:43:23 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-wm1-f74.google.com (mail-wm1-f74.google.com [209.85.128.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 1515668AB71 for ; Mon, 20 Dec 2021 16:43:17 +0200 (EET) Received: by mail-wm1-f74.google.com with SMTP id 85-20020a1c0158000000b003459d5d4867so1503721wmb.0 for ; Mon, 20 Dec 2021 06:43:17 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:in-reply-to:message-id:mime-version:references:subject:from:to :cc; bh=/Xq0HeS19EhUXMYpsSjK5RVWupaND+ZvALr+xEDlhNY=; b=SClSBDsaZexZM4vaZ26AAJKwiV/fWtB7jgxvM8tMm5INVa/z/My71Mx9zZpeDi7qsC rFi5q7j57Lxf0AkN6gs6gfPpFjx3vSECQzyzP48HDm9Y/IK9qp7Jl/aQGb7v9osSzquN vU6iGIwDQ7CW6kM/H8beeUWXLP2I7fHttyxmiMi0WMy0E7J4b6N69sWO5mHt9KIzX2TP HgNsXSN/SMzZce1C53ciGd8jRNunN9r21w4j5X0wiPMRZryufG9JAPyC/rhwSlGOg6om qgR9Bmv+d0eoAClUgM4RTpv+1cdiPnAt2qv7h5XISLUcvBS0NYn3oQ+4QpP+TKHKqXJF LpWA== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:in-reply-to:message-id:mime-version :references:subject:from:to:cc; bh=/Xq0HeS19EhUXMYpsSjK5RVWupaND+ZvALr+xEDlhNY=; b=Zkp/HGwNDMw6FYil4h2u6nX9S7wbsjIqaW08rxfYbiYANgXZvJNY/DzF+2mOgouw8+ CxL928Scfxk185Az0h5PLOCzocqhubNQVIIE3TT2BqNEUWC79YSjrR9M+0eFDAgU1Nm2 EIwW3mfl+zcNUbPVrCzmI0ychKpS5vKCOKqmwZb+/52xa81fAGLlIMgfcLxYoA8KhxWT dAl9TMb4+Ojv0eVGYM/ot/lROaFNYcEQZ5qE+0o2lYzwgn/MqIwIeO8WEJ1ujza9hr4I VeeFu355KYrI/VTJyIWhWDGdp88hw/0nOs6X12f6Vcz2u3QMqUSUOWUOJ3jrWrXnmpwr KYmw== X-Gm-Message-State: AOAM530hFHHb/FPE/JMLIeV75/GGliHSFszY6adKan6vENNbtUjV4Da/ aRSeGThugYwHwyvJQVng3MQFGU+FNFGkXS/E25etA8Bgw7KFjRysAjB5MTX6tRdSF37srvoNtwD g8SBH+pltWuekXmZd7gXXEdByMf9UM1WrKcObp55JUp3Q4EgzQAFZ67CGrXRe+xgVK/ubk0A= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:922d:7ddd:85f5:5a25]) (user=alankelly job=sendgmr) by 2002:a7b:c219:: with SMTP id x25mr10223wmi.1.1640011396131; Mon, 20 Dec 2021 06:43:16 -0800 (PST) Date: Mon, 20 Dec 2021 15:43:12 +0100 In-Reply-To: <8424d6a1-df63-954e-6823-740bf1fcb891@gmail.com> Message-Id: <20211220144312.738559-1-alankelly@google.com> Mime-Version: 1.0 References: <8424d6a1-df63-954e-6823-740bf1fcb891@gmail.com> X-Mailer: git-send-email 2.34.1.173.g76aa8bc2d0-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 1/2] libavutil/cpu: Add AV_CPU_FLAG_SLOW_GATHER. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: TTMcbMDSNK6f This flag is set on Haswell and earlier and all AMD cpus. --- Removes unnecessary indentation, clarifies comment and only sets flag on AMD cpus with AVX2. libavutil/cpu.h | 1 + libavutil/x86/cpu.c | 14 +++++++++++++- 2 files changed, 14 insertions(+), 1 deletion(-) diff --git a/libavutil/cpu.h b/libavutil/cpu.h index ae443eccad..ce9bf14bf7 100644 --- a/libavutil/cpu.h +++ b/libavutil/cpu.h @@ -54,6 +54,7 @@ #define AV_CPU_FLAG_BMI1 0x20000 ///< Bit Manipulation Instruction Set 1 #define AV_CPU_FLAG_BMI2 0x40000 ///< Bit Manipulation Instruction Set 2 #define AV_CPU_FLAG_AVX512 0x100000 ///< AVX-512 functions: requires OS support even if YMM/ZMM registers aren't used +#define AV_CPU_FLAG_SLOW_GATHER 0x2000000 ///< CPU has slow gathers. #define AV_CPU_FLAG_ALTIVEC 0x0001 ///< standard #define AV_CPU_FLAG_VSX 0x0002 ///< ISA 2.06 diff --git a/libavutil/x86/cpu.c b/libavutil/x86/cpu.c index bcd41a50a2..563984f234 100644 --- a/libavutil/x86/cpu.c +++ b/libavutil/x86/cpu.c @@ -146,8 +146,16 @@ int ff_get_cpu_flags_x86(void) if (max_std_level >= 7) { cpuid(7, eax, ebx, ecx, edx); #if HAVE_AVX2 - if ((rval & AV_CPU_FLAG_AVX) && (ebx & 0x00000020)) + if ((rval & AV_CPU_FLAG_AVX) && (ebx & 0x00000020)) { rval |= AV_CPU_FLAG_AVX2; + cpuid(1, eax, ebx, ecx, std_caps); + family = ((eax >> 8) & 0xf) + ((eax >> 20) & 0xff); + model = ((eax >> 4) & 0xf) + ((eax >> 12) & 0xf0); + /* Haswell has slow gather */ + if(family == 6 && model < 70) + rval |= AV_CPU_FLAG_SLOW_GATHER; + } + #if HAVE_AVX512 /* F, CD, BW, DQ, VL */ if ((xcr0_lo & 0xe0) == 0xe0) { /* OPMASK/ZMM state */ if ((rval & AV_CPU_FLAG_AVX2) && (ebx & 0xd0030000) == 0xd0030000) @@ -196,6 +204,10 @@ int ff_get_cpu_flags_x86(void) used unless explicitly disabled by checking AV_CPU_FLAG_AVXSLOW. */ if ((family == 0x15 || family == 0x16) && (rval & AV_CPU_FLAG_AVX)) rval |= AV_CPU_FLAG_AVXSLOW; + + /* AMD cpus have slow gather */ + if(rval & AV_CPU_FLAG_AVX2) + rval |= AV_CPU_FLAG_SLOW_GATHER; } /* XOP and FMA4 use the AVX instruction coding scheme, so they can't be From patchwork Mon Dec 20 14:45:45 2021 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Alan Kelly X-Patchwork-Id: 32757 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a6b:cd86:0:0:0:0:0 with SMTP id d128csp4429215iog; Mon, 20 Dec 2021 06:45:59 -0800 (PST) X-Google-Smtp-Source: ABdhPJzH3+zhc1gC8YqkbUWBrzh0lLR6VhfI0QhIi7Py3PYm/xFwbML2TFUbgeqXLM1wWADg+wV+ X-Received: by 2002:a17:906:1db2:: with SMTP id u18mr13285539ejh.729.1640011559069; Mon, 20 Dec 2021 06:45:59 -0800 (PST) ARC-Seal: i=1; a=rsa-sha256; t=1640011559; cv=none; d=google.com; s=arc-20160816; b=r4OZx0c0t2sdzd7bcxDwERGGp3IRTgDdqfnnbYQNadn7LsUOmbSwfefLD5M77OsBBI kdwplQqwOV3iO2qSJBtXRRQ8yy+IjwLRwwCzJYPXT/ZRd3HraYC9OnNal/iP1nmarcaN +y7CpG3PJlrvcI0X/sMi8MUmaS4SK6PdR2DnHiUQyi3s7GQYAU1r4DtA/ZSlI+LL3Cfs ZWQXy4n3RxLVB29xtl/QNR5B+3lBq7wgo/sjfMG+GMmWMirXDM10+YUvUFlugQElw1y7 C7QtTV232/rujaac+B+wyBFdMVJL5eJGjOw7iru6VNzW8Iw7aoqt0Omx9gK+cPIY84UT EC/g== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:cc:reply-to :list-subscribe:list-help:list-post:list-archive:list-unsubscribe :list-id:precedence:subject:to:from:references:mime-version :message-id:in-reply-to:date:dkim-signature:delivered-to; bh=usjJ3s6yz6QEH5ljZP0cW7QWW7lXVlaOcTc2WDjbesA=; b=iqGDblqoc3sf1PhIpVeUlZUs/zLPtGjSsmtmyoo9MMlBPMj3jiXbhycIZsNN2ENn2B HpeHQ0X+6wg104+aw2HfwDsC0Ivrwm2B+4F7bgO4fXLIQDBr4EJuTRNutLGe7SN6n67c 8hKyQn2I5CS1KUc1FTSTvOUNhA7BN8V8Vvs/C6o+FQ68UK642wfNW7C/bcPPazMxIm2F TiMJOChoIXFWfegBDj3zcqPcqdqd1vZ/DbOCjQocLTGodLsVpY7RTgpFhD53Y5pXKokX uwHBTKFLImFTcAxPnGgQ5+CZwGsTL1CxtADAuRM3k+9ngpUWW+kUj655GQH84aYAIpY6 uNzQ== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=iskcmjm4; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id x16si3326668ejo.967.2021.12.20.06.45.58; Mon, 20 Dec 2021 06:45:59 -0800 (PST) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@google.com header.s=20210112 header.b=iskcmjm4; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id CBD2C68AE86; Mon, 20 Dec 2021 16:45:55 +0200 (EET) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from mail-lf1-f74.google.com (mail-lf1-f74.google.com [209.85.167.74]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id 9E1FC68AB71 for ; Mon, 20 Dec 2021 16:45:49 +0200 (EET) Received: by mail-lf1-f74.google.com with SMTP id cf27-20020a056512281b00b004259e7fce67so1033206lfb.0 for ; Mon, 20 Dec 2021 06:45:49 -0800 (PST) DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=20210112; h=date:in-reply-to:message-id:mime-version:references:subject:from:to :cc; bh=2X2EepDowYoA+E+PZCbzwYE6q9udLErWYvTIa8HFSRI=; b=iskcmjm4e/YovR80hHlHiSMX4m86uBEu6fZjyscneRNecZVkGda5j1CyRwMnJ6uOoi LI0pDARejBSav0xjv2OGs4cq2d108xLr6g64halhoEbE5d5TR14DM3Vdt2SXnhrsDY3H n6oQ2F1uzhLGZiuN0dVJUmUyTVYGvlR6yDnoE/HF5pLgt8US183K6anRasNJ/rH3wySp tRv6M7t9HbWw7Cp5SVTIK3L0nuTfIeW04hOQj+ODzLw2uOqP7Ceq7wKNxz/VeQ+ZnydI Y1kORa1T9wjSYX/0xYPnitBFFTCu/yCJ3iMq7cxgEza85Fil3abzSCtRHkauUj/BVUmv ldGQ== X-Google-DKIM-Signature: v=1; a=rsa-sha256; c=relaxed/relaxed; d=1e100.net; s=20210112; h=x-gm-message-state:date:in-reply-to:message-id:mime-version :references:subject:from:to:cc; bh=2X2EepDowYoA+E+PZCbzwYE6q9udLErWYvTIa8HFSRI=; b=za5RRl/TdpFCy6i8+Emy6UhJhyxfkxB6qrnqT2Dw5pR51iCv3xP3fIfTaJkM/nesdC J6MjRtDXhQ0hpKoafDrE5fk73++IXrRKNVoPuULj9gQWFBPr6ndKED/GF/BN71oNjX8O LCU0kBIEf7awny/3EtkXvKlxjYScnouejc7A6/vh2ciX4gqNQmsrey8MVxtgRAlA41/J PZJIA/toFbVMBPvWi67fFhatLmkOylsgAVp6vYsr91VYyaY5yiWgBr3wvjReGS1uD3iZ 4c4Vw+UhsXiqwv63BCWbPTU40GDQbYrKd+7WuFNiUW0or5Rdqs93MqBERvX5HayNL5oD oQrQ== X-Gm-Message-State: AOAM533R32/52Gnsq585EL8jI1RU5+X1yTl105Bqgu9e0aQ+GtlOtyGa ZvGCTTmblN7AkeYvLriWI51bE2ZO0qm55osF2bXGgkf0ozfJ/fYKrHMEmnnXrglnhOT8EnKqM50 nUlV+SoGWwAkkd6H0Tfsp0lzDAi/aA9RtC9GUySQtR6KhvvYFKuTtniQY/kQ3RCrF7sX66gg= X-Received: from alankelly0.zrh.corp.google.com ([2a00:79e0:61:301:922d:7ddd:85f5:5a25]) (user=alankelly job=sendgmr) by 2002:a05:6512:2289:: with SMTP id f9mr15712665lfu.619.1640011548660; Mon, 20 Dec 2021 06:45:48 -0800 (PST) Date: Mon, 20 Dec 2021 15:45:45 +0100 In-Reply-To: <166a93ba-7f15-f473-5889-1a0a879a75a4@gmail.com> Message-Id: <20211220144545.739340-1-alankelly@google.com> Mime-Version: 1.0 References: <166a93ba-7f15-f473-5889-1a0a879a75a4@gmail.com> X-Mailer: git-send-email 2.34.1.173.g76aa8bc2d0-goog From: Alan Kelly To: ffmpeg-devel@ffmpeg.org Subject: [FFmpeg-devel] [PATCH 2/2] libswscale: Test AV_CPU_FLAG_SLOW_GATHER for hscale functions. X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Cc: Alan Kelly Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: x/bIDVReMEq+ This is instead of EXTERNAL_AVX2_FAST so that the avx2 hscale functions are only used where they are faster. --- Whoops! Corrects check so that this flag is only enabled where fast avx2 and fast gathers are available. libswscale/utils.c | 2 +- libswscale/x86/swscale.c | 2 +- tests/checkasm/sw_scale.c | 2 +- 3 files changed, 3 insertions(+), 3 deletions(-) diff --git a/libswscale/utils.c b/libswscale/utils.c index d4a72d3ce1..7158384f0b 100644 --- a/libswscale/utils.c +++ b/libswscale/utils.c @@ -282,7 +282,7 @@ void ff_shuffle_filter_coefficients(SwsContext *c, int *filterPos, int filterSiz #if ARCH_X86_64 int i, j, k, l; int cpu_flags = av_get_cpu_flags(); - if (EXTERNAL_AVX2_FAST(cpu_flags)){ + if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)){ if (dstW % 16 == 0){ if (filter != NULL){ diff --git a/libswscale/x86/swscale.c b/libswscale/x86/swscale.c index c49a05c37b..ffc7691c12 100644 --- a/libswscale/x86/swscale.c +++ b/libswscale/x86/swscale.c @@ -578,7 +578,7 @@ switch(c->dstBpc){ \ break; \ } - if (EXTERNAL_AVX2_FAST(cpu_flags)) { + if (EXTERNAL_AVX2_FAST(cpu_flags) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) { if ((c->srcBpc == 8) && (c->dstBpc <= 14)) { if (c->chrDstW % 16 == 0) ASSIGN_AVX2_SCALE_FUNC(c->hcScale, c->hChrFilterSize); diff --git a/tests/checkasm/sw_scale.c b/tests/checkasm/sw_scale.c index f4912e6c2c..3c0a083b42 100644 --- a/tests/checkasm/sw_scale.c +++ b/tests/checkasm/sw_scale.c @@ -217,7 +217,7 @@ static void check_hscale(void) } ff_sws_init_scale(ctx); memcpy(filterAvx2, filter, sizeof(uint16_t) * (SRC_PIXELS * MAX_FILTER_WIDTH + MAX_FILTER_WIDTH)); - if (cpu_flags & AV_CPU_FLAG_AVX2) + if ((cpu_flags & AV_CPU_FLAG_AVX2) && !(cpu_flags & AV_CPU_FLAG_SLOW_GATHER)) ff_shuffle_filter_coefficients(ctx, filterPosAvx, width, filterAvx2, SRC_PIXELS); if (check_func(ctx->hcScale, "hscale_%d_to_%d_width%d", ctx->srcBpc, ctx->dstBpc + 1, width)) {