From patchwork Wed Jun 15 19:51:24 2022 Content-Type: text/plain; charset="utf-8" MIME-Version: 1.0 Content-Transfer-Encoding: 7bit X-Patchwork-Submitter: Nil Admirari X-Patchwork-Id: 36243 Delivered-To: ffmpegpatchwork2@gmail.com Received: by 2002:a05:6a20:1a22:b0:84:42e0:ad30 with SMTP id cj34csp552654pzb; Wed, 15 Jun 2022 12:52:02 -0700 (PDT) X-Google-Smtp-Source: AGRyM1s/GGk2rSo1Za+uZM1vaL1YpzZ4ptmaC4eAZgCt1e4QYQNnRq4CdDA2Z1EGkeyCfygO5fyC X-Received: by 2002:a17:906:2f92:b0:711:cfe6:b117 with SMTP id w18-20020a1709062f9200b00711cfe6b117mr1225237eji.665.1655322722603; Wed, 15 Jun 2022 12:52:02 -0700 (PDT) ARC-Seal: i=1; a=rsa-sha256; t=1655322722; cv=none; d=google.com; s=arc-20160816; b=s4Z1A3Rvd0NWYDyfoSNt8Qh2N59TfVzsQ2DcGPNOdSykewS0Lu+tEXeOlLcuOeB6AZ VmxxA1EQ8SVYe+c3lDgCp9R46Jvc7g8IA2Yjqqi44qdkjAySFUD57FMe3CtDSqS0w5+C YvBND41hPiwIzy/pEF1xhFY88pl+5xCQL+2P89MgaVMS5QC6JAEMqgEXNOiCJmjSAZCD dIevAvGJn4L0D8XU34DnFsxNXcmXrEQahPXCjLHW+TAEOYstZrVKea1j4Z4FH3jpJH04 RI/KAEmceoGQKYGnkQe22jnsaC7XkE3d/Ubeyji+HBbeGdPyaiZRbWMxtQmM91Snm3+L 0enA== ARC-Message-Signature: i=1; a=rsa-sha256; c=relaxed/relaxed; d=google.com; s=arc-20160816; h=sender:errors-to:content-transfer-encoding:reply-to:list-subscribe :list-help:list-post:list-archive:list-unsubscribe:list-id :precedence:subject:mime-version:message-id:date:to:from :dkim-signature:delivered-to; bh=2ffIg+FEiDDFyaXhW6yzFlxt10PbAFuSYX9v1lMJkts=; b=oMwXlN0tdXHEzrcKfV0MX/NPrbKfJdQlLwxsRp10s8VhstsRI8QmzeZ8DHSxBWlCQl wqzeTyz6gW7MIYlUJQlBuNvvoARzV88sDnWvIJ0OkaKyNxh6Ly8sJA5hhmbYLYATppoD TF4ajBclQyMZema5XtGIAGXsW5+qHampIl1PgbfRSz8en1+XqFRWDHQPFAie+STT1RWE v+HLl1bMGYIbibgc10RW18Zbq5E4tzZLefJIHSDaObI/NxvBC78oPfhaw6UdllfPPaDk 5Z2A1H7xG9wsHIirEFJ+9rOSwuMz4946037ZJYCU146eo9pZ0YLhs8wtkCqqQRlXcvOU 62vg== ARC-Authentication-Results: i=1; mx.google.com; dkim=neutral (body hash did not verify) header.i=@mailo.com header.s=mailo header.b=LSbxu47T; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=mailo.com Return-Path: Received: from ffbox0-bg.mplayerhq.hu (ffbox0-bg.ffmpeg.org. [79.124.17.100]) by mx.google.com with ESMTP id w3-20020aa7da43000000b0043174e68c8esi71900eds.364.2022.06.15.12.52.02; Wed, 15 Jun 2022 12:52:02 -0700 (PDT) Received-SPF: pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) client-ip=79.124.17.100; Authentication-Results: mx.google.com; dkim=neutral (body hash did not verify) header.i=@mailo.com header.s=mailo header.b=LSbxu47T; spf=pass (google.com: domain of ffmpeg-devel-bounces@ffmpeg.org designates 79.124.17.100 as permitted sender) smtp.mailfrom=ffmpeg-devel-bounces@ffmpeg.org; dmarc=fail (p=NONE sp=NONE dis=NONE) header.from=mailo.com Received: from [127.0.1.1] (localhost [127.0.0.1]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTP id 20A3468B6F0; Wed, 15 Jun 2022 22:51:48 +0300 (EEST) X-Original-To: ffmpeg-devel@ffmpeg.org Delivered-To: ffmpeg-devel@ffmpeg.org Received: from msg-1.mailo.com (msg-1.mailo.com [213.182.54.11]) by ffbox0-bg.mplayerhq.hu (Postfix) with ESMTPS id D5E8E68B02F for ; Wed, 15 Jun 2022 22:51:40 +0300 (EEST) DKIM-Signature: v=1; a=rsa-sha256; c=simple/simple; d=mailo.com; s=mailo; t=1655322695; bh=qM65tz5uT1neNwCdUeAIP3WcUO2oL5skPRv/6VbMdR0=; h=X-EA-Auth:From:To:Subject:Date:Message-Id:X-Mailer:MIME-Version: Content-Transfer-Encoding; b=LSbxu47THORRrsx0NEmcwuK54MqUk3zhgGBU3RGIghXYfp9ELEXKqWn/0othMi8bC zEJP1cdpJpGG3lQeqK898MzFzM4rn+eBLPMZOhAdbYFfx9YbGa9bY3lShl4n52wTBi +bLhAVx3QPz6HlTxUztd+w2ICSm3fZKuLgPeTDPE= Received: by b-4.in.mailobj.net [192.168.90.14] with ESMTP via ip-206.mailobj.net [213.182.55.206] Wed, 15 Jun 2022 21:51:35 +0200 (CEST) X-EA-Auth: oGOAh2YrQXqUA4d7lb+4qgg3BKUoyNvJXn+wi29hL70bfvjwDtjt/T7Ye62of2uRq28keckDYCagMOm8K7rA+Nvunm9sHpcl+De6KHrYllk= From: Nil Admirari To: ffmpeg-devel@ffmpeg.org Date: Wed, 15 Jun 2022 22:51:24 +0300 Message-Id: <20220615195128.15796-1-nil-admirari@mailo.com> X-Mailer: git-send-email 2.34.1 MIME-Version: 1.0 Subject: [FFmpeg-devel] [PATCH v15 1/5] libavutil: Add wchartoutf8(), wchartoansi(), utf8toansi() and getenv_utf8() X-BeenThere: ffmpeg-devel@ffmpeg.org X-Mailman-Version: 2.1.29 Precedence: list List-Id: FFmpeg development discussions and patches List-Unsubscribe: , List-Archive: List-Post: List-Help: List-Subscribe: , Reply-To: FFmpeg development discussions and patches Errors-To: ffmpeg-devel-bounces@ffmpeg.org Sender: "ffmpeg-devel" X-TUID: uLs7z0pUEcQE wchartoutf8() converts strings returned by WinAPI into UTF-8, which is FFmpeg's preffered encoding. Some external dependencies, such as AviSynth, are still not Unicode-enabled. utf8toansi() converts UTF-8 strings into ANSI in two steps: UTF-8 -> wchar_t -> ANSI. wchartoansi() is responsible for the second step of the conversion. Conversion in just one step is not supported by WinAPI. Since these character converting functions allocate the buffer of necessary size, they also facilitate the removal of MAX_PATH limit in places where fixed-size ANSI/WCHAR strings were used as filename buffers. getenv_utf8() wraps _wgetenv() converting its input from and its output to UTF-8. Compared to plain getenv(), getenv_utf8() requires a cleanup. Because of that, in places that only test the existence of an environment variable or compare its value with a string consisting entirely of ASCII characters, the use of plain getenv() is still preferred. (libavutil/log.c check_color_terminal() is an example of such a place.) Plain getenv() is also preffered in UNIX-only code, such as bktr.c, fbdev_common.c, oss.c in libavdevice or af_ladspa.c in libavfilter. --- .vscode/settings.json | 11 ++++++ configure | 1 + libavutil/getenv_utf8.h | 71 ++++++++++++++++++++++++++++++++++++++ libavutil/wchar_filename.h | 51 +++++++++++++++++++++++++++ 4 files changed, 134 insertions(+) create mode 100644 .vscode/settings.json create mode 100644 libavutil/getenv_utf8.h diff --git a/.vscode/settings.json b/.vscode/settings.json new file mode 100644 index 0000000000..e866d57743 --- /dev/null +++ b/.vscode/settings.json @@ -0,0 +1,11 @@ +{ + "files.associations": { + "w32dlfcn.h": "c", + "mem.h": "c", + "internal.h": "c", + "os_support.h": "c", + "packet_internal.h": "c", + "stdlib.h": "c", + "mathematics.h": "c" + } +} \ No newline at end of file diff --git a/configure b/configure index 3dca1c4bd3..fa37a74531 100755 --- a/configure +++ b/configure @@ -2272,6 +2272,7 @@ SYSTEM_FUNCS=" fcntl getaddrinfo getauxval + getenv gethrtime getopt GetModuleHandle diff --git a/libavutil/getenv_utf8.h b/libavutil/getenv_utf8.h new file mode 100644 index 0000000000..161e3e6202 --- /dev/null +++ b/libavutil/getenv_utf8.h @@ -0,0 +1,71 @@ +/* + * This file is part of FFmpeg. + * + * FFmpeg is free software; you can redistribute it and/or + * modify it under the terms of the GNU Lesser General Public + * License as published by the Free Software Foundation; either + * version 2.1 of the License, or (at your option) any later version. + * + * FFmpeg is distributed in the hope that it will be useful, + * but WITHOUT ANY WARRANTY; without even the implied warranty of + * MERCHANTABILITY or FITNESS FOR A PARTICULAR PURPOSE. See the GNU + * Lesser General Public License for more details. + * + * You should have received a copy of the GNU Lesser General Public + * License along with FFmpeg; if not, write to the Free Software + * Foundation, Inc., 51 Franklin Street, Fifth Floor, Boston, MA 02110-1301 USA + */ + +#ifndef AVUTIL_GETENV_UTF8_H +#define AVUTIL_GETENV_UTF8_H + +#include + +#include "mem.h" + +#ifdef HAVE_GETENV + +#ifdef _WIN32 + +#include "libavutil/wchar_filename.h" + +static inline char *getenv_utf8(const char *varname) +{ + wchar_t *varname_w, *var_w; + char *var; + + if (utf8towchar(varname, &varname_w)) + return NULL; + if (!varname_w) + return NULL; + + var_w = _wgetenv(varname_w); + av_free(varname_w); + + if (!var_w) + return NULL; + if (wchartoutf8(var_w, &var)) + return NULL; + + return var; + + // No CP_ACP fallback compared to other *_utf8() functions: + // non UTF-8 strings must not be returned. +} + +#else + +static inline char *getenv_utf8(const char *varname) +{ + return av_strdup(getenv(varname)); +} + +#endif // _WIN32 + +#else + +#define getenv_utf8(x) NULL + +#endif // HAVE_GETENV + +#endif // AVUTIL_GETENV_UTF8_H diff --git a/libavutil/wchar_filename.h b/libavutil/wchar_filename.h index f36d9dfea3..a6d71e52e5 100644 --- a/libavutil/wchar_filename.h +++ b/libavutil/wchar_filename.h @@ -41,6 +41,57 @@ static inline int utf8towchar(const char *filename_utf8, wchar_t **filename_w) return 0; } +av_warn_unused_result +static inline int wchartocp(unsigned int code_page, const wchar_t *filename_w, + char **filename) +{ + DWORD flags = code_page == CP_UTF8 ? WC_ERR_INVALID_CHARS : 0; + int num_chars = WideCharToMultiByte(code_page, flags, filename_w, -1, + NULL, 0, NULL, NULL); + if (num_chars <= 0) { + *filename = NULL; + return 0; + } + *filename = av_malloc_array(num_chars, sizeof *filename); + if (!*filename) { + errno = ENOMEM; + return -1; + } + WideCharToMultiByte(code_page, flags, filename_w, -1, + *filename, num_chars, NULL, NULL); + return 0; +} + +av_warn_unused_result +static inline int wchartoutf8(const wchar_t *filename_w, char **filename) +{ + return wchartocp(CP_UTF8, filename_w, filename); +} + +av_warn_unused_result +static inline int wchartoansi(const wchar_t *filename_w, char **filename) +{ + return wchartocp(CP_ACP, filename_w, filename); +} + +av_warn_unused_result +static inline int utf8toansi(const char *filename_utf8, char **filename) +{ + wchar_t *filename_w = NULL; + int ret = -1; + if (utf8towchar(filename_utf8, &filename_w)) + return -1; + + if (!filename_w) { + *filename = NULL; + return 0; + } + + ret = wchartoansi(filename_w, filename); + av_free(filename_w); + return ret; +} + /** * Checks for extended path prefixes for which normalization needs to be skipped. * see .NET6: PathInternal.IsExtended()