<!DOCTYPE html> <html lang="en" data-content_root="../../../"> <head> <meta charset="utf-8" /> <meta name="viewport" content="width=device-width, initial-scale=1.0" /> <meta name="viewport" content="width=device-width, initial-scale=1"> <title>searx.engines.startpage — SearXNG Documentation (2025.5.28+2288f07d6)</title> <link rel="stylesheet" type="text/css" href="../../../_static/pygments.css?v=6625fa76" /> <link rel="stylesheet" type="text/css" href="../../../_static/searxng.css?v=52e4ff28" /> <script src="../../../_static/documentation_options.js?v=9a91f0a1"></script> <script src="../../../_static/doctools.js?v=9a2dae69"></script> <script src="../../../_static/sphinx_highlight.js?v=dc90522c"></script> <script data-project="searxng" data-version="2025.5.28+2288f07d6" src="../../../_static/describe_version.js?v=fa7f30d0"></script> <link rel="index" title="Index" href="../../../genindex.html" /> <link rel="search" title="Search" href="../../../search.html" /> </head><body> <div class="related" role="navigation" aria-label="Related"> <h3>Navigation</h3> <ul> <li class="right" style="margin-right: 10px"> <a href="../../../genindex.html" title="General Index" accesskey="I">index</a></li> <li class="right" > <a href="../../../py-modindex.html" title="Python Module Index" >modules</a> |</li> <li class="nav-item nav-item-0"><a href="../../../index.html">SearXNG Documentation (2025.5.28+2288f07d6)</a> »</li> <li class="nav-item nav-item-1"><a href="../../index.html" >Module code</a> »</li> <li class="nav-item nav-item-2"><a href="../engines.html" accesskey="U">searx.engines</a> »</li> <li class="nav-item nav-item-this"><a href="">searx.engines.startpage</a></li> </ul> </div> <div class="document"> <div class="documentwrapper"> <div class="bodywrapper"> <div class="body" role="main"> <h1>Source code for searx.engines.startpage</h1><div class="highlight"><pre> <span></span><span class="c1"># SPDX-License-Identifier: AGPL-3.0-or-later</span> <span class="sd">"""Startpage's language & region selectors are a mess ..</span> <span class="sd">.. _startpage regions:</span> <span class="sd">Startpage regions</span> <span class="sd">=================</span> <span class="sd">In the list of regions there are tags we need to map to common region tags::</span> <span class="sd"> pt-BR_BR --> pt_BR</span> <span class="sd"> zh-CN_CN --> zh_Hans_CN</span> <span class="sd"> zh-TW_TW --> zh_Hant_TW</span> <span class="sd"> zh-TW_HK --> zh_Hant_HK</span> <span class="sd"> en-GB_GB --> en_GB</span> <span class="sd">and there is at least one tag with a three letter language tag (ISO 639-2)::</span> <span class="sd"> fil_PH --> fil_PH</span> <span class="sd">The locale code ``no_NO`` from Startpage does not exists and is mapped to</span> <span class="sd">``nb-NO``::</span> <span class="sd"> babel.core.UnknownLocaleError: unknown locale 'no_NO'</span> <span class="sd">For reference see languages-subtag at iana; ``no`` is the macrolanguage [1]_ and</span> <span class="sd">W3C recommends subtag over macrolanguage [2]_.</span> <span class="sd">.. [1] `iana: language-subtag-registry</span> <span class="sd"> <https://www.iana.org/assignments/language-subtag-registry/language-subtag-registry>`_ ::</span> <span class="sd"> type: language</span> <span class="sd"> Subtag: nb</span> <span class="sd"> Description: Norwegian Bokmål</span> <span class="sd"> Added: 2005-10-16</span> <span class="sd"> Suppress-Script: Latn</span> <span class="sd"> Macrolanguage: no</span> <span class="sd">.. [2]</span> <span class="sd"> Use macrolanguages with care. Some language subtags have a Scope field set to</span> <span class="sd"> macrolanguage, i.e. this primary language subtag encompasses a number of more</span> <span class="sd"> specific primary language subtags in the registry. ... As we recommended for</span> <span class="sd"> the collection subtags mentioned above, in most cases you should try to use</span> <span class="sd"> the more specific subtags ... `W3: The primary language subtag</span> <span class="sd"> <https://www.w3.org/International/questions/qa-choosing-language-tags#langsubtag>`_</span> <span class="sd">.. _startpage languages:</span> <span class="sd">Startpage languages</span> <span class="sd">===================</span> <span class="sd">:py:obj:`send_accept_language_header`:</span> <span class="sd"> The displayed name in Startpage's settings page depend on the location of the</span> <span class="sd"> IP when ``Accept-Language`` HTTP header is unset. In :py:obj:`fetch_traits`</span> <span class="sd"> we use::</span> <span class="sd"> 'Accept-Language': "en-US,en;q=0.5",</span> <span class="sd"> ..</span> <span class="sd"> to get uniform names independent from the IP).</span> <span class="sd">.. _startpage categories:</span> <span class="sd">Startpage categories</span> <span class="sd">====================</span> <span class="sd">Startpage's category (for Web-search, News, Videos, ..) is set by</span> <span class="sd">:py:obj:`startpage_categ` in settings.yml::</span> <span class="sd"> - name: startpage</span> <span class="sd"> engine: startpage</span> <span class="sd"> startpage_categ: web</span> <span class="sd"> ...</span> <span class="sd">.. hint::</span> <span class="sd"> Supported categories are ``web``, ``news`` and ``images``.</span> <span class="sd">"""</span> <span class="c1"># pylint: disable=too-many-statements</span> <span class="kn">from</span><span class="w"> </span><span class="nn">__future__</span><span class="w"> </span><span class="kn">import</span> <span class="n">annotations</span> <span class="kn">from</span><span class="w"> </span><span class="nn">typing</span><span class="w"> </span><span class="kn">import</span> <span class="n">TYPE_CHECKING</span><span class="p">,</span> <span class="n">Any</span> <span class="kn">from</span><span class="w"> </span><span class="nn">collections</span><span class="w"> </span><span class="kn">import</span> <span class="n">OrderedDict</span> <span class="kn">import</span><span class="w"> </span><span class="nn">re</span> <span class="kn">from</span><span class="w"> </span><span class="nn">unicodedata</span><span class="w"> </span><span class="kn">import</span> <span class="n">normalize</span><span class="p">,</span> <span class="n">combining</span> <span class="kn">from</span><span class="w"> </span><span class="nn">datetime</span><span class="w"> </span><span class="kn">import</span> <span class="n">datetime</span><span class="p">,</span> <span class="n">timedelta</span> <span class="kn">from</span><span class="w"> </span><span class="nn">json</span><span class="w"> </span><span class="kn">import</span> <span class="n">loads</span> <span class="kn">import</span><span class="w"> </span><span class="nn">dateutil.parser</span> <span class="kn">import</span><span class="w"> </span><span class="nn">lxml.html</span> <span class="kn">import</span><span class="w"> </span><span class="nn">babel.localedata</span> <span class="kn">from</span><span class="w"> </span><span class="nn">searx.utils</span><span class="w"> </span><span class="kn">import</span> <span class="n">extr</span><span class="p">,</span> <span class="n">extract_text</span><span class="p">,</span> <span class="n">eval_xpath</span><span class="p">,</span> <span class="n">gen_useragent</span><span class="p">,</span> <span class="n">html_to_text</span><span class="p">,</span> <span class="n">humanize_bytes</span><span class="p">,</span> <span class="n">remove_pua_from_str</span> <span class="kn">from</span><span class="w"> </span><span class="nn">searx.network</span><span class="w"> </span><span class="kn">import</span> <span class="n">get</span> <span class="c1"># see https://github.com/searxng/searxng/issues/762</span> <span class="kn">from</span><span class="w"> </span><span class="nn">searx.exceptions</span><span class="w"> </span><span class="kn">import</span> <span class="n">SearxEngineCaptchaException</span> <span class="kn">from</span><span class="w"> </span><span class="nn">searx.locales</span><span class="w"> </span><span class="kn">import</span> <span class="n">region_tag</span> <span class="kn">from</span><span class="w"> </span><span class="nn">searx.enginelib.traits</span><span class="w"> </span><span class="kn">import</span> <span class="n">EngineTraits</span> <span class="kn">from</span><span class="w"> </span><span class="nn">searx.enginelib</span><span class="w"> </span><span class="kn">import</span> <span class="n">EngineCache</span> <span class="k">if</span> <span class="n">TYPE_CHECKING</span><span class="p">:</span> <span class="kn">import</span><span class="w"> </span><span class="nn">logging</span> <span class="n">logger</span><span class="p">:</span> <span class="n">logging</span><span class="o">.</span><span class="n">Logger</span> <span class="n">traits</span><span class="p">:</span> <span class="n">EngineTraits</span> <span class="c1"># about</span> <span class="n">about</span> <span class="o">=</span> <span class="p">{</span> <span class="s2">"website"</span><span class="p">:</span> <span class="s1">'https://startpage.com'</span><span class="p">,</span> <span class="s2">"wikidata_id"</span><span class="p">:</span> <span class="s1">'Q2333295'</span><span class="p">,</span> <span class="s2">"official_api_documentation"</span><span class="p">:</span> <span class="kc">None</span><span class="p">,</span> <span class="s2">"use_official_api"</span><span class="p">:</span> <span class="kc">False</span><span class="p">,</span> <span class="s2">"require_api_key"</span><span class="p">:</span> <span class="kc">False</span><span class="p">,</span> <span class="s2">"results"</span><span class="p">:</span> <span class="s1">'HTML'</span><span class="p">,</span> <span class="p">}</span> <span class="n">startpage_categ</span> <span class="o">=</span> <span class="s1">'web'</span> <span class="sd">"""Startpage's category, visit :ref:`startpage categories`.</span> <span class="sd">"""</span> <span class="n">send_accept_language_header</span> <span class="o">=</span> <span class="kc">True</span> <span class="sd">"""Startpage tries to guess user's language and territory from the HTTP</span> <span class="sd">``Accept-Language``. Optional the user can select a search-language (can be</span> <span class="sd">different to the UI language) and a region filter.</span> <span class="sd">"""</span> <span class="c1"># engine dependent config</span> <span class="n">categories</span> <span class="o">=</span> <span class="p">[</span><span class="s1">'general'</span><span class="p">,</span> <span class="s1">'web'</span><span class="p">]</span> <span class="n">paging</span> <span class="o">=</span> <span class="kc">True</span> <span class="n">max_page</span> <span class="o">=</span> <span class="mi">18</span> <span class="sd">"""Tested 18 pages maximum (argument ``page``), to be save max is set to 20."""</span> <span class="n">time_range_support</span> <span class="o">=</span> <span class="kc">True</span> <span class="n">safesearch</span> <span class="o">=</span> <span class="kc">True</span> <span class="n">time_range_dict</span> <span class="o">=</span> <span class="p">{</span><span class="s1">'day'</span><span class="p">:</span> <span class="s1">'d'</span><span class="p">,</span> <span class="s1">'week'</span><span class="p">:</span> <span class="s1">'w'</span><span class="p">,</span> <span class="s1">'month'</span><span class="p">:</span> <span class="s1">'m'</span><span class="p">,</span> <span class="s1">'year'</span><span class="p">:</span> <span class="s1">'y'</span><span class="p">}</span> <span class="n">safesearch_dict</span> <span class="o">=</span> <span class="p">{</span><span class="mi">0</span><span class="p">:</span> <span class="s1">'0'</span><span class="p">,</span> <span class="mi">1</span><span class="p">:</span> <span class="s1">'1'</span><span class="p">,</span> <span class="mi">2</span><span class="p">:</span> <span class="s1">'1'</span><span class="p">}</span> <span class="c1"># search-url</span> <span class="n">base_url</span> <span class="o">=</span> <span class="s1">'https://www.startpage.com'</span> <span class="n">search_url</span> <span class="o">=</span> <span class="n">base_url</span> <span class="o">+</span> <span class="s1">'/sp/search'</span> <span class="c1"># specific xpath variables</span> <span class="c1"># ads xpath //div[@id="results"]/div[@id="sponsored"]//div[@class="result"]</span> <span class="c1"># not ads: div[@class="result"] are the direct children of div[@id="results"]</span> <span class="n">search_form_xpath</span> <span class="o">=</span> <span class="s1">'//form[@id="search"]'</span> <span class="sd">"""XPath of Startpage's origin search form</span> <span class="sd">.. code: html</span> <span class="sd"> <form action="/sp/search" method="post"></span> <span class="sd"> <input type="text" name="query" value="" ..></span> <span class="sd"> <input type="hidden" name="t" value="device"></span> <span class="sd"> <input type="hidden" name="lui" value="english"></span> <span class="sd"> <input type="hidden" name="sc" value="Q7Mt5TRqowKB00"></span> <span class="sd"> <input type="hidden" name="cat" value="web"></span> <span class="sd"> <input type="hidden" class="abp" id="abp-input" name="abp" value="1"></span> <span class="sd"> </form></span> <span class="sd">"""</span> <span class="n">CACHE</span><span class="p">:</span> <span class="n">EngineCache</span> <span class="sd">"""Persistent (SQLite) key/value cache that deletes its values after ``expire``</span> <span class="sd">seconds."""</span> <span class="k">def</span><span class="w"> </span><span class="nf">init</span><span class="p">(</span><span class="n">_</span><span class="p">):</span> <span class="k">global</span> <span class="n">CACHE</span> <span class="c1"># pylint: disable=global-statement</span> <span class="c1"># hint: all three startpage engines (WEB, Images & News) can/should use the</span> <span class="c1"># same sc_code ..</span> <span class="n">CACHE</span> <span class="o">=</span> <span class="n">EngineCache</span><span class="p">(</span><span class="s2">"startpage"</span><span class="p">)</span> <span class="c1"># type:ignore</span> <span class="n">sc_code_cache_sec</span> <span class="o">=</span> <span class="mi">3600</span> <span class="sd">"""Time in seconds the sc-code is cached in memory :py:obj:`get_sc_code`."""</span> <div class="viewcode-block" id="get_sc_code"> <a class="viewcode-back" href="../../../dev/engines/online/startpage.html#searx.engines.startpage.get_sc_code">[docs]</a> <span class="k">def</span><span class="w"> </span><span class="nf">get_sc_code</span><span class="p">(</span><span class="n">searxng_locale</span><span class="p">,</span> <span class="n">params</span><span class="p">):</span> <span class="w"> </span><span class="sd">"""Get an actual ``sc`` argument from Startpage's search form (HTML page).</span> <span class="sd"> Startpage puts a ``sc`` argument on every HTML :py:obj:`search form</span> <span class="sd"> <search_form_xpath>`. Without this argument Startpage considers the request</span> <span class="sd"> is from a bot. We do not know what is encoded in the value of the ``sc``</span> <span class="sd"> argument, but it seems to be a kind of a *time-stamp*.</span> <span class="sd"> Startpage's search form generates a new sc-code on each request. This</span> <span class="sd"> function scrap a new sc-code from Startpage's home page every</span> <span class="sd"> :py:obj:`sc_code_cache_sec` seconds."""</span> <span class="n">sc_code</span> <span class="o">=</span> <span class="n">CACHE</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s2">"SC_CODE"</span><span class="p">,</span> <span class="s2">""</span><span class="p">)</span> <span class="k">if</span> <span class="n">sc_code</span><span class="p">:</span> <span class="k">return</span> <span class="n">sc_code</span> <span class="n">headers</span> <span class="o">=</span> <span class="p">{</span><span class="o">**</span><span class="n">params</span><span class="p">[</span><span class="s1">'headers'</span><span class="p">]}</span> <span class="n">headers</span><span class="p">[</span><span class="s1">'Origin'</span><span class="p">]</span> <span class="o">=</span> <span class="n">base_url</span> <span class="n">headers</span><span class="p">[</span><span class="s1">'Referer'</span><span class="p">]</span> <span class="o">=</span> <span class="n">base_url</span> <span class="o">+</span> <span class="s1">'/'</span> <span class="c1"># headers['Connection'] = 'keep-alive'</span> <span class="c1"># headers['Accept-Encoding'] = 'gzip, deflate, br'</span> <span class="c1"># headers['Accept'] = 'text/html,application/xhtml+xml,application/xml;q=0.9,image/avif,image/webp,*/*;q=0.8'</span> <span class="c1"># headers['User-Agent'] = 'Mozilla/5.0 (X11; Ubuntu; Linux x86_64; rv:105.0) Gecko/20100101 Firefox/105.0'</span> <span class="c1"># add Accept-Language header</span> <span class="k">if</span> <span class="n">searxng_locale</span> <span class="o">==</span> <span class="s1">'all'</span><span class="p">:</span> <span class="n">searxng_locale</span> <span class="o">=</span> <span class="s1">'en-US'</span> <span class="n">locale</span> <span class="o">=</span> <span class="n">babel</span><span class="o">.</span><span class="n">Locale</span><span class="o">.</span><span class="n">parse</span><span class="p">(</span><span class="n">searxng_locale</span><span class="p">,</span> <span class="n">sep</span><span class="o">=</span><span class="s1">'-'</span><span class="p">)</span> <span class="k">if</span> <span class="n">send_accept_language_header</span><span class="p">:</span> <span class="n">ac_lang</span> <span class="o">=</span> <span class="n">locale</span><span class="o">.</span><span class="n">language</span> <span class="k">if</span> <span class="n">locale</span><span class="o">.</span><span class="n">territory</span><span class="p">:</span> <span class="n">ac_lang</span> <span class="o">=</span> <span class="s2">"</span><span class="si">%s</span><span class="s2">-</span><span class="si">%s</span><span class="s2">,</span><span class="si">%s</span><span class="s2">;q=0.9,*;q=0.5"</span> <span class="o">%</span> <span class="p">(</span> <span class="n">locale</span><span class="o">.</span><span class="n">language</span><span class="p">,</span> <span class="n">locale</span><span class="o">.</span><span class="n">territory</span><span class="p">,</span> <span class="n">locale</span><span class="o">.</span><span class="n">language</span><span class="p">,</span> <span class="p">)</span> <span class="n">headers</span><span class="p">[</span><span class="s1">'Accept-Language'</span><span class="p">]</span> <span class="o">=</span> <span class="n">ac_lang</span> <span class="n">get_sc_url</span> <span class="o">=</span> <span class="n">base_url</span> <span class="o">+</span> <span class="s1">'/?sc=</span><span class="si">%s</span><span class="s1">'</span> <span class="o">%</span> <span class="p">(</span><span class="n">sc_code</span><span class="p">)</span> <span class="n">logger</span><span class="o">.</span><span class="n">debug</span><span class="p">(</span><span class="s2">"query new sc time-stamp ... </span><span class="si">%s</span><span class="s2">"</span><span class="p">,</span> <span class="n">get_sc_url</span><span class="p">)</span> <span class="n">logger</span><span class="o">.</span><span class="n">debug</span><span class="p">(</span><span class="s2">"headers: </span><span class="si">%s</span><span class="s2">"</span><span class="p">,</span> <span class="n">headers</span><span class="p">)</span> <span class="n">resp</span> <span class="o">=</span> <span class="n">get</span><span class="p">(</span><span class="n">get_sc_url</span><span class="p">,</span> <span class="n">headers</span><span class="o">=</span><span class="n">headers</span><span class="p">)</span> <span class="c1"># ?? x = network.get('https://www.startpage.com/sp/cdn/images/filter-chevron.svg', headers=headers)</span> <span class="c1"># ?? https://www.startpage.com/sp/cdn/images/filter-chevron.svg</span> <span class="c1"># ?? ping-back URL: https://www.startpage.com/sp/pb?sc=TLsB0oITjZ8F21</span> <span class="k">if</span> <span class="nb">str</span><span class="p">(</span><span class="n">resp</span><span class="o">.</span><span class="n">url</span><span class="p">)</span><span class="o">.</span><span class="n">startswith</span><span class="p">(</span><span class="s1">'https://www.startpage.com/sp/captcha'</span><span class="p">):</span> <span class="c1"># type: ignore</span> <span class="k">raise</span> <span class="n">SearxEngineCaptchaException</span><span class="p">(</span> <span class="n">message</span><span class="o">=</span><span class="s2">"get_sc_code: got redirected to https://www.startpage.com/sp/captcha"</span><span class="p">,</span> <span class="p">)</span> <span class="n">dom</span> <span class="o">=</span> <span class="n">lxml</span><span class="o">.</span><span class="n">html</span><span class="o">.</span><span class="n">fromstring</span><span class="p">(</span><span class="n">resp</span><span class="o">.</span><span class="n">text</span><span class="p">)</span> <span class="c1"># type: ignore</span> <span class="k">try</span><span class="p">:</span> <span class="n">sc_code</span> <span class="o">=</span> <span class="n">eval_xpath</span><span class="p">(</span><span class="n">dom</span><span class="p">,</span> <span class="n">search_form_xpath</span> <span class="o">+</span> <span class="s1">'//input[@name="sc"]/@value'</span><span class="p">)[</span><span class="mi">0</span><span class="p">]</span> <span class="k">except</span> <span class="ne">IndexError</span> <span class="k">as</span> <span class="n">exc</span><span class="p">:</span> <span class="n">logger</span><span class="o">.</span><span class="n">debug</span><span class="p">(</span><span class="s2">"suspend startpage API --> https://github.com/searxng/searxng/pull/695"</span><span class="p">)</span> <span class="k">raise</span> <span class="n">SearxEngineCaptchaException</span><span class="p">(</span> <span class="n">message</span><span class="o">=</span><span class="s2">"get_sc_code: [PR-695] query new sc time-stamp failed! (</span><span class="si">%s</span><span class="s2">)"</span> <span class="o">%</span> <span class="n">resp</span><span class="o">.</span><span class="n">url</span><span class="p">,</span> <span class="c1"># type: ignore</span> <span class="p">)</span> <span class="kn">from</span><span class="w"> </span><span class="nn">exc</span> <span class="n">sc_code</span> <span class="o">=</span> <span class="nb">str</span><span class="p">(</span><span class="n">sc_code</span><span class="p">)</span> <span class="n">logger</span><span class="o">.</span><span class="n">debug</span><span class="p">(</span><span class="s2">"get_sc_code: new value is: </span><span class="si">%s</span><span class="s2">"</span><span class="p">,</span> <span class="n">sc_code</span><span class="p">)</span> <span class="n">CACHE</span><span class="o">.</span><span class="n">set</span><span class="p">(</span><span class="n">key</span><span class="o">=</span><span class="s2">"SC_CODE"</span><span class="p">,</span> <span class="n">value</span><span class="o">=</span><span class="n">sc_code</span><span class="p">,</span> <span class="n">expire</span><span class="o">=</span><span class="n">sc_code_cache_sec</span><span class="p">)</span> <span class="k">return</span> <span class="n">sc_code</span></div> <div class="viewcode-block" id="request"> <a class="viewcode-back" href="../../../dev/engines/online/startpage.html#searx.engines.startpage.request">[docs]</a> <span class="k">def</span><span class="w"> </span><span class="nf">request</span><span class="p">(</span><span class="n">query</span><span class="p">,</span> <span class="n">params</span><span class="p">):</span> <span class="w"> </span><span class="sd">"""Assemble a Startpage request.</span> <span class="sd"> To avoid CAPTCHA we need to send a well formed HTTP POST request with a</span> <span class="sd"> cookie. We need to form a request that is identical to the request build by</span> <span class="sd"> Startpage's search form:</span> <span class="sd"> - in the cookie the **region** is selected</span> <span class="sd"> - in the HTTP POST data the **language** is selected</span> <span class="sd"> Additionally the arguments form Startpage's search form needs to be set in</span> <span class="sd"> HTML POST data / compare ``<input>`` elements: :py:obj:`search_form_xpath`.</span> <span class="sd"> """</span> <span class="n">engine_region</span> <span class="o">=</span> <span class="n">traits</span><span class="o">.</span><span class="n">get_region</span><span class="p">(</span><span class="n">params</span><span class="p">[</span><span class="s1">'searxng_locale'</span><span class="p">],</span> <span class="s1">'en-US'</span><span class="p">)</span> <span class="n">engine_language</span> <span class="o">=</span> <span class="n">traits</span><span class="o">.</span><span class="n">get_language</span><span class="p">(</span><span class="n">params</span><span class="p">[</span><span class="s1">'searxng_locale'</span><span class="p">],</span> <span class="s1">'en'</span><span class="p">)</span> <span class="c1"># build arguments</span> <span class="n">args</span> <span class="o">=</span> <span class="p">{</span> <span class="s1">'query'</span><span class="p">:</span> <span class="n">query</span><span class="p">,</span> <span class="s1">'cat'</span><span class="p">:</span> <span class="n">startpage_categ</span><span class="p">,</span> <span class="s1">'t'</span><span class="p">:</span> <span class="s1">'device'</span><span class="p">,</span> <span class="s1">'sc'</span><span class="p">:</span> <span class="n">get_sc_code</span><span class="p">(</span><span class="n">params</span><span class="p">[</span><span class="s1">'searxng_locale'</span><span class="p">],</span> <span class="n">params</span><span class="p">),</span> <span class="c1"># hint: this func needs HTTP headers,</span> <span class="s1">'with_date'</span><span class="p">:</span> <span class="n">time_range_dict</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">params</span><span class="p">[</span><span class="s1">'time_range'</span><span class="p">],</span> <span class="s1">''</span><span class="p">),</span> <span class="p">}</span> <span class="k">if</span> <span class="n">engine_language</span><span class="p">:</span> <span class="n">args</span><span class="p">[</span><span class="s1">'language'</span><span class="p">]</span> <span class="o">=</span> <span class="n">engine_language</span> <span class="n">args</span><span class="p">[</span><span class="s1">'lui'</span><span class="p">]</span> <span class="o">=</span> <span class="n">engine_language</span> <span class="n">args</span><span class="p">[</span><span class="s1">'abp'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'1'</span> <span class="k">if</span> <span class="n">params</span><span class="p">[</span><span class="s1">'pageno'</span><span class="p">]</span> <span class="o">></span> <span class="mi">1</span><span class="p">:</span> <span class="n">args</span><span class="p">[</span><span class="s1">'page'</span><span class="p">]</span> <span class="o">=</span> <span class="n">params</span><span class="p">[</span><span class="s1">'pageno'</span><span class="p">]</span> <span class="c1"># build cookie</span> <span class="n">lang_homepage</span> <span class="o">=</span> <span class="s1">'en'</span> <span class="n">cookie</span> <span class="o">=</span> <span class="n">OrderedDict</span><span class="p">()</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'date_time'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'world'</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'disable_family_filter'</span><span class="p">]</span> <span class="o">=</span> <span class="n">safesearch_dict</span><span class="p">[</span><span class="n">params</span><span class="p">[</span><span class="s1">'safesearch'</span><span class="p">]]</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'disable_open_in_new_window'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'0'</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'enable_post_method'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'1'</span> <span class="c1"># hint: POST</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'enable_proxy_safety_suggest'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'1'</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'enable_stay_control'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'1'</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'instant_answers'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'1'</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'lang_homepage'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'s/device/</span><span class="si">%s</span><span class="s1">/'</span> <span class="o">%</span> <span class="n">lang_homepage</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'num_of_results'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'10'</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'suggestions'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'1'</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'wt_unit'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'celsius'</span> <span class="k">if</span> <span class="n">engine_language</span><span class="p">:</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'language'</span><span class="p">]</span> <span class="o">=</span> <span class="n">engine_language</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'language_ui'</span><span class="p">]</span> <span class="o">=</span> <span class="n">engine_language</span> <span class="k">if</span> <span class="n">engine_region</span><span class="p">:</span> <span class="n">cookie</span><span class="p">[</span><span class="s1">'search_results_region'</span><span class="p">]</span> <span class="o">=</span> <span class="n">engine_region</span> <span class="n">params</span><span class="p">[</span><span class="s1">'cookies'</span><span class="p">][</span><span class="s1">'preferences'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'N1N'</span><span class="o">.</span><span class="n">join</span><span class="p">([</span><span class="s2">"</span><span class="si">%s</span><span class="s2">EEE</span><span class="si">%s</span><span class="s2">"</span> <span class="o">%</span> <span class="n">x</span> <span class="k">for</span> <span class="n">x</span> <span class="ow">in</span> <span class="n">cookie</span><span class="o">.</span><span class="n">items</span><span class="p">()])</span> <span class="n">logger</span><span class="o">.</span><span class="n">debug</span><span class="p">(</span><span class="s1">'cookie preferences: </span><span class="si">%s</span><span class="s1">'</span><span class="p">,</span> <span class="n">params</span><span class="p">[</span><span class="s1">'cookies'</span><span class="p">][</span><span class="s1">'preferences'</span><span class="p">])</span> <span class="c1"># POST request</span> <span class="n">logger</span><span class="o">.</span><span class="n">debug</span><span class="p">(</span><span class="s2">"data: </span><span class="si">%s</span><span class="s2">"</span><span class="p">,</span> <span class="n">args</span><span class="p">)</span> <span class="n">params</span><span class="p">[</span><span class="s1">'data'</span><span class="p">]</span> <span class="o">=</span> <span class="n">args</span> <span class="n">params</span><span class="p">[</span><span class="s1">'method'</span><span class="p">]</span> <span class="o">=</span> <span class="s1">'POST'</span> <span class="n">params</span><span class="p">[</span><span class="s1">'url'</span><span class="p">]</span> <span class="o">=</span> <span class="n">search_url</span> <span class="n">params</span><span class="p">[</span><span class="s1">'headers'</span><span class="p">][</span><span class="s1">'Origin'</span><span class="p">]</span> <span class="o">=</span> <span class="n">base_url</span> <span class="n">params</span><span class="p">[</span><span class="s1">'headers'</span><span class="p">][</span><span class="s1">'Referer'</span><span class="p">]</span> <span class="o">=</span> <span class="n">base_url</span> <span class="o">+</span> <span class="s1">'/'</span> <span class="c1"># is the Accept header needed?</span> <span class="c1"># params['headers']['Accept'] = 'text/html,application/xhtml+xml,application/xml;q=0.9,image/webp,*/*;q=0.8'</span> <span class="k">return</span> <span class="n">params</span></div> <span class="k">def</span><span class="w"> </span><span class="nf">_parse_published_date</span><span class="p">(</span><span class="n">content</span><span class="p">:</span> <span class="nb">str</span><span class="p">)</span> <span class="o">-></span> <span class="nb">tuple</span><span class="p">[</span><span class="nb">str</span><span class="p">,</span> <span class="n">datetime</span> <span class="o">|</span> <span class="kc">None</span><span class="p">]:</span> <span class="n">published_date</span> <span class="o">=</span> <span class="kc">None</span> <span class="c1"># check if search result starts with something like: "2 Sep 2014 ... "</span> <span class="k">if</span> <span class="n">re</span><span class="o">.</span><span class="n">match</span><span class="p">(</span><span class="sa">r</span><span class="s2">"^([1-9]|[1-2][0-9]|3[0-1]) [A-Z][a-z]</span><span class="si">{2}</span><span class="s2"> [0-9]</span><span class="si">{4}</span><span class="s2"> \.\.\. "</span><span class="p">,</span> <span class="n">content</span><span class="p">):</span> <span class="n">date_pos</span> <span class="o">=</span> <span class="n">content</span><span class="o">.</span><span class="n">find</span><span class="p">(</span><span class="s1">'...'</span><span class="p">)</span> <span class="o">+</span> <span class="mi">4</span> <span class="n">date_string</span> <span class="o">=</span> <span class="n">content</span><span class="p">[</span><span class="mi">0</span> <span class="p">:</span> <span class="n">date_pos</span> <span class="o">-</span> <span class="mi">5</span><span class="p">]</span> <span class="c1"># fix content string</span> <span class="n">content</span> <span class="o">=</span> <span class="n">content</span><span class="p">[</span><span class="n">date_pos</span><span class="p">:]</span> <span class="k">try</span><span class="p">:</span> <span class="n">published_date</span> <span class="o">=</span> <span class="n">dateutil</span><span class="o">.</span><span class="n">parser</span><span class="o">.</span><span class="n">parse</span><span class="p">(</span><span class="n">date_string</span><span class="p">,</span> <span class="n">dayfirst</span><span class="o">=</span><span class="kc">True</span><span class="p">)</span> <span class="k">except</span> <span class="ne">ValueError</span><span class="p">:</span> <span class="k">pass</span> <span class="c1"># check if search result starts with something like: "5 days ago ... "</span> <span class="k">elif</span> <span class="n">re</span><span class="o">.</span><span class="n">match</span><span class="p">(</span><span class="sa">r</span><span class="s2">"^[0-9]+ days? ago \.\.\. "</span><span class="p">,</span> <span class="n">content</span><span class="p">):</span> <span class="n">date_pos</span> <span class="o">=</span> <span class="n">content</span><span class="o">.</span><span class="n">find</span><span class="p">(</span><span class="s1">'...'</span><span class="p">)</span> <span class="o">+</span> <span class="mi">4</span> <span class="n">date_string</span> <span class="o">=</span> <span class="n">content</span><span class="p">[</span><span class="mi">0</span> <span class="p">:</span> <span class="n">date_pos</span> <span class="o">-</span> <span class="mi">5</span><span class="p">]</span> <span class="c1"># calculate datetime</span> <span class="n">published_date</span> <span class="o">=</span> <span class="n">datetime</span><span class="o">.</span><span class="n">now</span><span class="p">()</span> <span class="o">-</span> <span class="n">timedelta</span><span class="p">(</span><span class="n">days</span><span class="o">=</span><span class="nb">int</span><span class="p">(</span><span class="n">re</span><span class="o">.</span><span class="n">match</span><span class="p">(</span><span class="sa">r</span><span class="s1">'\d+'</span><span class="p">,</span> <span class="n">date_string</span><span class="p">)</span><span class="o">.</span><span class="n">group</span><span class="p">()))</span> <span class="c1"># type: ignore</span> <span class="c1"># fix content string</span> <span class="n">content</span> <span class="o">=</span> <span class="n">content</span><span class="p">[</span><span class="n">date_pos</span><span class="p">:]</span> <span class="k">return</span> <span class="n">content</span><span class="p">,</span> <span class="n">published_date</span> <span class="k">def</span><span class="w"> </span><span class="nf">_get_web_result</span><span class="p">(</span><span class="n">result</span><span class="p">):</span> <span class="n">content</span> <span class="o">=</span> <span class="n">html_to_text</span><span class="p">(</span><span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'description'</span><span class="p">))</span> <span class="n">content</span><span class="p">,</span> <span class="n">publishedDate</span> <span class="o">=</span> <span class="n">_parse_published_date</span><span class="p">(</span><span class="n">content</span><span class="p">)</span> <span class="k">return</span> <span class="p">{</span> <span class="s1">'url'</span><span class="p">:</span> <span class="n">result</span><span class="p">[</span><span class="s1">'clickUrl'</span><span class="p">],</span> <span class="s1">'title'</span><span class="p">:</span> <span class="n">html_to_text</span><span class="p">(</span><span class="n">result</span><span class="p">[</span><span class="s1">'title'</span><span class="p">]),</span> <span class="s1">'content'</span><span class="p">:</span> <span class="n">content</span><span class="p">,</span> <span class="s1">'publishedDate'</span><span class="p">:</span> <span class="n">publishedDate</span><span class="p">,</span> <span class="p">}</span> <span class="k">def</span><span class="w"> </span><span class="nf">_get_news_result</span><span class="p">(</span><span class="n">result</span><span class="p">):</span> <span class="n">title</span> <span class="o">=</span> <span class="n">remove_pua_from_str</span><span class="p">(</span><span class="n">html_to_text</span><span class="p">(</span><span class="n">result</span><span class="p">[</span><span class="s1">'title'</span><span class="p">]))</span> <span class="n">content</span> <span class="o">=</span> <span class="n">remove_pua_from_str</span><span class="p">(</span><span class="n">html_to_text</span><span class="p">(</span><span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'description'</span><span class="p">)))</span> <span class="n">publishedDate</span> <span class="o">=</span> <span class="kc">None</span> <span class="k">if</span> <span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'date'</span><span class="p">):</span> <span class="n">publishedDate</span> <span class="o">=</span> <span class="n">datetime</span><span class="o">.</span><span class="n">fromtimestamp</span><span class="p">(</span><span class="n">result</span><span class="p">[</span><span class="s1">'date'</span><span class="p">]</span> <span class="o">/</span> <span class="mi">1000</span><span class="p">)</span> <span class="n">thumbnailUrl</span> <span class="o">=</span> <span class="kc">None</span> <span class="k">if</span> <span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'thumbnailUrl'</span><span class="p">):</span> <span class="n">thumbnailUrl</span> <span class="o">=</span> <span class="n">base_url</span> <span class="o">+</span> <span class="n">result</span><span class="p">[</span><span class="s1">'thumbnailUrl'</span><span class="p">]</span> <span class="k">return</span> <span class="p">{</span> <span class="s1">'url'</span><span class="p">:</span> <span class="n">result</span><span class="p">[</span><span class="s1">'clickUrl'</span><span class="p">],</span> <span class="s1">'title'</span><span class="p">:</span> <span class="n">title</span><span class="p">,</span> <span class="s1">'content'</span><span class="p">:</span> <span class="n">content</span><span class="p">,</span> <span class="s1">'publishedDate'</span><span class="p">:</span> <span class="n">publishedDate</span><span class="p">,</span> <span class="s1">'thumbnail'</span><span class="p">:</span> <span class="n">thumbnailUrl</span><span class="p">,</span> <span class="p">}</span> <span class="k">def</span><span class="w"> </span><span class="nf">_get_image_result</span><span class="p">(</span><span class="n">result</span><span class="p">)</span> <span class="o">-></span> <span class="nb">dict</span><span class="p">[</span><span class="nb">str</span><span class="p">,</span> <span class="n">Any</span><span class="p">]</span> <span class="o">|</span> <span class="kc">None</span><span class="p">:</span> <span class="n">url</span> <span class="o">=</span> <span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'altClickUrl'</span><span class="p">)</span> <span class="k">if</span> <span class="ow">not</span> <span class="n">url</span><span class="p">:</span> <span class="k">return</span> <span class="kc">None</span> <span class="n">thumbnailUrl</span> <span class="o">=</span> <span class="kc">None</span> <span class="k">if</span> <span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'thumbnailUrl'</span><span class="p">):</span> <span class="n">thumbnailUrl</span> <span class="o">=</span> <span class="n">base_url</span> <span class="o">+</span> <span class="n">result</span><span class="p">[</span><span class="s1">'thumbnailUrl'</span><span class="p">]</span> <span class="n">resolution</span> <span class="o">=</span> <span class="kc">None</span> <span class="k">if</span> <span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'width'</span><span class="p">)</span> <span class="ow">and</span> <span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'height'</span><span class="p">):</span> <span class="n">resolution</span> <span class="o">=</span> <span class="sa">f</span><span class="s2">"</span><span class="si">{</span><span class="n">result</span><span class="p">[</span><span class="s1">'width'</span><span class="p">]</span><span class="si">}</span><span class="s2">x</span><span class="si">{</span><span class="n">result</span><span class="p">[</span><span class="s1">'height'</span><span class="p">]</span><span class="si">}</span><span class="s2">"</span> <span class="n">filesize</span> <span class="o">=</span> <span class="kc">None</span> <span class="k">if</span> <span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'filesize'</span><span class="p">):</span> <span class="n">size_str</span> <span class="o">=</span> <span class="s1">''</span><span class="o">.</span><span class="n">join</span><span class="p">(</span><span class="nb">filter</span><span class="p">(</span><span class="nb">str</span><span class="o">.</span><span class="n">isdigit</span><span class="p">,</span> <span class="n">result</span><span class="p">[</span><span class="s1">'filesize'</span><span class="p">]))</span> <span class="n">filesize</span> <span class="o">=</span> <span class="n">humanize_bytes</span><span class="p">(</span><span class="nb">int</span><span class="p">(</span><span class="n">size_str</span><span class="p">))</span> <span class="k">return</span> <span class="p">{</span> <span class="s1">'template'</span><span class="p">:</span> <span class="s1">'images.html'</span><span class="p">,</span> <span class="s1">'url'</span><span class="p">:</span> <span class="n">url</span><span class="p">,</span> <span class="s1">'title'</span><span class="p">:</span> <span class="n">html_to_text</span><span class="p">(</span><span class="n">result</span><span class="p">[</span><span class="s1">'title'</span><span class="p">]),</span> <span class="s1">'content'</span><span class="p">:</span> <span class="s1">''</span><span class="p">,</span> <span class="s1">'img_src'</span><span class="p">:</span> <span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'rawImageUrl'</span><span class="p">),</span> <span class="s1">'thumbnail_src'</span><span class="p">:</span> <span class="n">thumbnailUrl</span><span class="p">,</span> <span class="s1">'resolution'</span><span class="p">:</span> <span class="n">resolution</span><span class="p">,</span> <span class="s1">'img_format'</span><span class="p">:</span> <span class="n">result</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'format'</span><span class="p">),</span> <span class="s1">'filesize'</span><span class="p">:</span> <span class="n">filesize</span><span class="p">,</span> <span class="p">}</span> <span class="k">def</span><span class="w"> </span><span class="nf">response</span><span class="p">(</span><span class="n">resp</span><span class="p">):</span> <span class="n">categ</span> <span class="o">=</span> <span class="n">startpage_categ</span><span class="o">.</span><span class="n">capitalize</span><span class="p">()</span> <span class="n">results_raw</span> <span class="o">=</span> <span class="s1">'{'</span> <span class="o">+</span> <span class="n">extr</span><span class="p">(</span><span class="n">resp</span><span class="o">.</span><span class="n">text</span><span class="p">,</span> <span class="sa">f</span><span class="s2">"React.createElement(UIStartpage.AppSerp</span><span class="si">{</span><span class="n">categ</span><span class="si">}</span><span class="s2">, </span><span class="se">{{</span><span class="s2">"</span><span class="p">,</span> <span class="s1">'}})'</span><span class="p">)</span> <span class="o">+</span> <span class="s1">'}}'</span> <span class="n">results_json</span> <span class="o">=</span> <span class="n">loads</span><span class="p">(</span><span class="n">results_raw</span><span class="p">)</span> <span class="n">results_obj</span> <span class="o">=</span> <span class="n">results_json</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'render'</span><span class="p">,</span> <span class="p">{})</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'presenter'</span><span class="p">,</span> <span class="p">{})</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'regions'</span><span class="p">,</span> <span class="p">{})</span> <span class="n">results</span> <span class="o">=</span> <span class="p">[]</span> <span class="k">for</span> <span class="n">results_categ</span> <span class="ow">in</span> <span class="n">results_obj</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'mainline'</span><span class="p">,</span> <span class="p">[]):</span> <span class="k">for</span> <span class="n">item</span> <span class="ow">in</span> <span class="n">results_categ</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'results'</span><span class="p">,</span> <span class="p">[]):</span> <span class="k">if</span> <span class="n">results_categ</span><span class="p">[</span><span class="s1">'display_type'</span><span class="p">]</span> <span class="o">==</span> <span class="s1">'web-google'</span><span class="p">:</span> <span class="n">results</span><span class="o">.</span><span class="n">append</span><span class="p">(</span><span class="n">_get_web_result</span><span class="p">(</span><span class="n">item</span><span class="p">))</span> <span class="k">elif</span> <span class="n">results_categ</span><span class="p">[</span><span class="s1">'display_type'</span><span class="p">]</span> <span class="o">==</span> <span class="s1">'news-bing'</span><span class="p">:</span> <span class="n">results</span><span class="o">.</span><span class="n">append</span><span class="p">(</span><span class="n">_get_news_result</span><span class="p">(</span><span class="n">item</span><span class="p">))</span> <span class="k">elif</span> <span class="s1">'images'</span> <span class="ow">in</span> <span class="n">results_categ</span><span class="p">[</span><span class="s1">'display_type'</span><span class="p">]:</span> <span class="n">item</span> <span class="o">=</span> <span class="n">_get_image_result</span><span class="p">(</span><span class="n">item</span><span class="p">)</span> <span class="k">if</span> <span class="n">item</span><span class="p">:</span> <span class="n">results</span><span class="o">.</span><span class="n">append</span><span class="p">(</span><span class="n">item</span><span class="p">)</span> <span class="k">return</span> <span class="n">results</span> <div class="viewcode-block" id="fetch_traits"> <a class="viewcode-back" href="../../../dev/engines/online/startpage.html#searx.engines.startpage.fetch_traits">[docs]</a> <span class="k">def</span><span class="w"> </span><span class="nf">fetch_traits</span><span class="p">(</span><span class="n">engine_traits</span><span class="p">:</span> <span class="n">EngineTraits</span><span class="p">):</span> <span class="w"> </span><span class="sd">"""Fetch :ref:`languages <startpage languages>` and :ref:`regions <startpage</span> <span class="sd"> regions>` from Startpage."""</span> <span class="c1"># pylint: disable=too-many-branches</span> <span class="n">headers</span> <span class="o">=</span> <span class="p">{</span> <span class="s1">'User-Agent'</span><span class="p">:</span> <span class="n">gen_useragent</span><span class="p">(),</span> <span class="s1">'Accept-Language'</span><span class="p">:</span> <span class="s2">"en-US,en;q=0.5"</span><span class="p">,</span> <span class="c1"># bing needs to set the English language</span> <span class="p">}</span> <span class="n">resp</span> <span class="o">=</span> <span class="n">get</span><span class="p">(</span><span class="s1">'https://www.startpage.com/do/settings'</span><span class="p">,</span> <span class="n">headers</span><span class="o">=</span><span class="n">headers</span><span class="p">)</span> <span class="k">if</span> <span class="ow">not</span> <span class="n">resp</span><span class="o">.</span><span class="n">ok</span><span class="p">:</span> <span class="c1"># type: ignore</span> <span class="nb">print</span><span class="p">(</span><span class="s2">"ERROR: response from Startpage is not OK."</span><span class="p">)</span> <span class="n">dom</span> <span class="o">=</span> <span class="n">lxml</span><span class="o">.</span><span class="n">html</span><span class="o">.</span><span class="n">fromstring</span><span class="p">(</span><span class="n">resp</span><span class="o">.</span><span class="n">text</span><span class="p">)</span> <span class="c1"># type: ignore</span> <span class="c1"># regions</span> <span class="n">sp_region_names</span> <span class="o">=</span> <span class="p">[]</span> <span class="k">for</span> <span class="n">option</span> <span class="ow">in</span> <span class="n">dom</span><span class="o">.</span><span class="n">xpath</span><span class="p">(</span><span class="s1">'//form[@name="settings"]//select[@name="search_results_region"]/option'</span><span class="p">):</span> <span class="n">sp_region_names</span><span class="o">.</span><span class="n">append</span><span class="p">(</span><span class="n">option</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'value'</span><span class="p">))</span> <span class="k">for</span> <span class="n">eng_tag</span> <span class="ow">in</span> <span class="n">sp_region_names</span><span class="p">:</span> <span class="k">if</span> <span class="n">eng_tag</span> <span class="o">==</span> <span class="s1">'all'</span><span class="p">:</span> <span class="k">continue</span> <span class="n">babel_region_tag</span> <span class="o">=</span> <span class="p">{</span><span class="s1">'no_NO'</span><span class="p">:</span> <span class="s1">'nb_NO'</span><span class="p">}</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">eng_tag</span><span class="p">,</span> <span class="n">eng_tag</span><span class="p">)</span> <span class="c1"># norway</span> <span class="k">if</span> <span class="s1">'-'</span> <span class="ow">in</span> <span class="n">babel_region_tag</span><span class="p">:</span> <span class="n">l</span><span class="p">,</span> <span class="n">r</span> <span class="o">=</span> <span class="n">babel_region_tag</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="s1">'-'</span><span class="p">)</span> <span class="n">r</span> <span class="o">=</span> <span class="n">r</span><span class="o">.</span><span class="n">split</span><span class="p">(</span><span class="s1">'_'</span><span class="p">)[</span><span class="o">-</span><span class="mi">1</span><span class="p">]</span> <span class="n">sxng_tag</span> <span class="o">=</span> <span class="n">region_tag</span><span class="p">(</span><span class="n">babel</span><span class="o">.</span><span class="n">Locale</span><span class="o">.</span><span class="n">parse</span><span class="p">(</span><span class="n">l</span> <span class="o">+</span> <span class="s1">'_'</span> <span class="o">+</span> <span class="n">r</span><span class="p">,</span> <span class="n">sep</span><span class="o">=</span><span class="s1">'_'</span><span class="p">))</span> <span class="k">else</span><span class="p">:</span> <span class="k">try</span><span class="p">:</span> <span class="n">sxng_tag</span> <span class="o">=</span> <span class="n">region_tag</span><span class="p">(</span><span class="n">babel</span><span class="o">.</span><span class="n">Locale</span><span class="o">.</span><span class="n">parse</span><span class="p">(</span><span class="n">babel_region_tag</span><span class="p">,</span> <span class="n">sep</span><span class="o">=</span><span class="s1">'_'</span><span class="p">))</span> <span class="k">except</span> <span class="n">babel</span><span class="o">.</span><span class="n">UnknownLocaleError</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="s2">"ERROR: can't determine babel locale of startpage's locale </span><span class="si">%s</span><span class="s2">"</span> <span class="o">%</span> <span class="n">eng_tag</span><span class="p">)</span> <span class="k">continue</span> <span class="n">conflict</span> <span class="o">=</span> <span class="n">engine_traits</span><span class="o">.</span><span class="n">regions</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">sxng_tag</span><span class="p">)</span> <span class="k">if</span> <span class="n">conflict</span><span class="p">:</span> <span class="k">if</span> <span class="n">conflict</span> <span class="o">!=</span> <span class="n">eng_tag</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="s2">"CONFLICT: babel </span><span class="si">%s</span><span class="s2"> --> </span><span class="si">%s</span><span class="s2">, </span><span class="si">%s</span><span class="s2">"</span> <span class="o">%</span> <span class="p">(</span><span class="n">sxng_tag</span><span class="p">,</span> <span class="n">conflict</span><span class="p">,</span> <span class="n">eng_tag</span><span class="p">))</span> <span class="k">continue</span> <span class="n">engine_traits</span><span class="o">.</span><span class="n">regions</span><span class="p">[</span><span class="n">sxng_tag</span><span class="p">]</span> <span class="o">=</span> <span class="n">eng_tag</span> <span class="c1"># languages</span> <span class="n">catalog_engine2code</span> <span class="o">=</span> <span class="p">{</span><span class="n">name</span><span class="o">.</span><span class="n">lower</span><span class="p">():</span> <span class="n">lang_code</span> <span class="k">for</span> <span class="n">lang_code</span><span class="p">,</span> <span class="n">name</span> <span class="ow">in</span> <span class="n">babel</span><span class="o">.</span><span class="n">Locale</span><span class="p">(</span><span class="s1">'en'</span><span class="p">)</span><span class="o">.</span><span class="n">languages</span><span class="o">.</span><span class="n">items</span><span class="p">()}</span> <span class="c1"># get the native name of every language known by babel</span> <span class="k">for</span> <span class="n">lang_code</span> <span class="ow">in</span> <span class="nb">filter</span><span class="p">(</span><span class="k">lambda</span> <span class="n">lang_code</span><span class="p">:</span> <span class="n">lang_code</span><span class="o">.</span><span class="n">find</span><span class="p">(</span><span class="s1">'_'</span><span class="p">)</span> <span class="o">==</span> <span class="o">-</span><span class="mi">1</span><span class="p">,</span> <span class="n">babel</span><span class="o">.</span><span class="n">localedata</span><span class="o">.</span><span class="n">locale_identifiers</span><span class="p">()):</span> <span class="n">native_name</span> <span class="o">=</span> <span class="n">babel</span><span class="o">.</span><span class="n">Locale</span><span class="p">(</span><span class="n">lang_code</span><span class="p">)</span><span class="o">.</span><span class="n">get_language_name</span><span class="p">()</span> <span class="k">if</span> <span class="ow">not</span> <span class="n">native_name</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="sa">f</span><span class="s2">"ERROR: language name of startpage's language </span><span class="si">{</span><span class="n">lang_code</span><span class="si">}</span><span class="s2"> is unknown by babel"</span><span class="p">)</span> <span class="k">continue</span> <span class="n">native_name</span> <span class="o">=</span> <span class="n">native_name</span><span class="o">.</span><span class="n">lower</span><span class="p">()</span> <span class="c1"># add native name exactly as it is</span> <span class="n">catalog_engine2code</span><span class="p">[</span><span class="n">native_name</span><span class="p">]</span> <span class="o">=</span> <span class="n">lang_code</span> <span class="c1"># add "normalized" language name (i.e. français becomes francais and español becomes espanol)</span> <span class="n">unaccented_name</span> <span class="o">=</span> <span class="s1">''</span><span class="o">.</span><span class="n">join</span><span class="p">(</span><span class="nb">filter</span><span class="p">(</span><span class="k">lambda</span> <span class="n">c</span><span class="p">:</span> <span class="ow">not</span> <span class="n">combining</span><span class="p">(</span><span class="n">c</span><span class="p">),</span> <span class="n">normalize</span><span class="p">(</span><span class="s1">'NFKD'</span><span class="p">,</span> <span class="n">native_name</span><span class="p">)))</span> <span class="k">if</span> <span class="nb">len</span><span class="p">(</span><span class="n">unaccented_name</span><span class="p">)</span> <span class="o">==</span> <span class="nb">len</span><span class="p">(</span><span class="n">unaccented_name</span><span class="o">.</span><span class="n">encode</span><span class="p">()):</span> <span class="c1"># add only if result is ascii (otherwise "normalization" didn't work)</span> <span class="n">catalog_engine2code</span><span class="p">[</span><span class="n">unaccented_name</span><span class="p">]</span> <span class="o">=</span> <span class="n">lang_code</span> <span class="c1"># values that can't be determined by babel's languages names</span> <span class="n">catalog_engine2code</span><span class="o">.</span><span class="n">update</span><span class="p">(</span> <span class="p">{</span> <span class="c1"># traditional chinese used in ..</span> <span class="s1">'fantizhengwen'</span><span class="p">:</span> <span class="s1">'zh_Hant'</span><span class="p">,</span> <span class="c1"># Korean alphabet</span> <span class="s1">'hangul'</span><span class="p">:</span> <span class="s1">'ko'</span><span class="p">,</span> <span class="c1"># Malayalam is one of 22 scheduled languages of India.</span> <span class="s1">'malayam'</span><span class="p">:</span> <span class="s1">'ml'</span><span class="p">,</span> <span class="s1">'norsk'</span><span class="p">:</span> <span class="s1">'nb'</span><span class="p">,</span> <span class="s1">'sinhalese'</span><span class="p">:</span> <span class="s1">'si'</span><span class="p">,</span> <span class="p">}</span> <span class="p">)</span> <span class="n">skip_eng_tags</span> <span class="o">=</span> <span class="p">{</span> <span class="s1">'english_uk'</span><span class="p">,</span> <span class="c1"># SearXNG lang 'en' already maps to 'english'</span> <span class="p">}</span> <span class="k">for</span> <span class="n">option</span> <span class="ow">in</span> <span class="n">dom</span><span class="o">.</span><span class="n">xpath</span><span class="p">(</span><span class="s1">'//form[@name="settings"]//select[@name="language"]/option'</span><span class="p">):</span> <span class="n">eng_tag</span> <span class="o">=</span> <span class="n">option</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="s1">'value'</span><span class="p">)</span> <span class="k">if</span> <span class="n">eng_tag</span> <span class="ow">in</span> <span class="n">skip_eng_tags</span><span class="p">:</span> <span class="k">continue</span> <span class="n">name</span> <span class="o">=</span> <span class="n">extract_text</span><span class="p">(</span><span class="n">option</span><span class="p">)</span><span class="o">.</span><span class="n">lower</span><span class="p">()</span> <span class="c1"># type: ignore</span> <span class="n">sxng_tag</span> <span class="o">=</span> <span class="n">catalog_engine2code</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">eng_tag</span><span class="p">)</span> <span class="k">if</span> <span class="n">sxng_tag</span> <span class="ow">is</span> <span class="kc">None</span><span class="p">:</span> <span class="n">sxng_tag</span> <span class="o">=</span> <span class="n">catalog_engine2code</span><span class="p">[</span><span class="n">name</span><span class="p">]</span> <span class="n">conflict</span> <span class="o">=</span> <span class="n">engine_traits</span><span class="o">.</span><span class="n">languages</span><span class="o">.</span><span class="n">get</span><span class="p">(</span><span class="n">sxng_tag</span><span class="p">)</span> <span class="k">if</span> <span class="n">conflict</span><span class="p">:</span> <span class="k">if</span> <span class="n">conflict</span> <span class="o">!=</span> <span class="n">eng_tag</span><span class="p">:</span> <span class="nb">print</span><span class="p">(</span><span class="s2">"CONFLICT: babel </span><span class="si">%s</span><span class="s2"> --> </span><span class="si">%s</span><span class="s2">, </span><span class="si">%s</span><span class="s2">"</span> <span class="o">%</span> <span class="p">(</span><span class="n">sxng_tag</span><span class="p">,</span> <span class="n">conflict</span><span class="p">,</span> <span class="n">eng_tag</span><span class="p">))</span> <span class="k">continue</span> <span class="n">engine_traits</span><span class="o">.</span><span class="n">languages</span><span class="p">[</span><span class="n">sxng_tag</span><span class="p">]</span> <span class="o">=</span> <span class="n">eng_tag</span></div> </pre></div> <div class="clearer"></div> </div> </div> </div> <span id="sidebar-top"></span> <div class="sphinxsidebar" role="navigation" aria-label="Main"> <div class="sphinxsidebarwrapper"> <p class="logo"><a href="../../../index.html"> <img class="logo" src="../../../_static/searxng-wordmark.svg" alt="Logo of SearXNG"/> </a></p> <h3><a href="../../../index.html">Table of Contents</a></h3> <ul> <li class="toctree-l1"><a class="reference internal" href="../../../user/index.html">User information</a></li> <li class="toctree-l1"><a class="reference internal" href="../../../own-instance.html">Why use a private instance?</a></li> <li class="toctree-l1"><a class="reference internal" href="../../../admin/index.html">Administrator documentation</a></li> <li class="toctree-l1"><a class="reference internal" href="../../../dev/index.html">Developer documentation</a></li> <li class="toctree-l1"><a class="reference internal" href="../../../utils/index.html">DevOps tooling box</a></li> <li class="toctree-l1"><a class="reference internal" href="../../../src/index.html">Source-Code</a></li> </ul> <h3>Project Links</h3> <ul> <li><a href="https://github.com/searxng/searxng/tree/master">Source</a> <li><a href="https://github.com/searxng/searxng/wiki">Wiki</a> <li><a href="https://searx.space">Public instances</a> <li><a href="https://github.com/searxng/searxng/issues">Issue Tracker</a> </ul><h3>Navigation</h3> <ul> <li><a href="../../../index.html">Overview</a> <ul> <li><a href="../../index.html">Module code</a> <ul> <li><a href="../engines.html">searx.engines</a> </ul> </li></ul> </li> </ul> </li> </ul> <search id="searchbox" style="display: none" role="search"> <h3 id="searchlabel">Quick search</h3> <div class="searchformwrapper"> <form class="search" action="../../../search.html" method="get"> <input type="text" name="q" aria-labelledby="searchlabel" autocomplete="off" autocorrect="off" autocapitalize="off" spellcheck="false"/> <input type="submit" value="Go" /> </form> </div> </search> <script>document.getElementById('searchbox').style.display = "block"</script> </div> </div> <div class="clearer"></div> </div> <div class="footer" role="contentinfo"> © Copyright SearXNG team. </div> </body> </html>