Skip to content

[Bug]: ridiculous limits on MP-API queries #1103

Description

@computron

Code snippet

"""Minimal repro: exclude_elements query string >60 chars triggers 422."""
from mp_api.client import MPRester

# 23 radioactive elements -> comma-joined string is 72 chars, exceeds API's 60-char cap
RADIOACTIVE = [
    "Tc", "Pm", "Po", "At", "Rn", "Fr", "Ra",
    "Ac", "Th", "Pa", "U", "Np", "Pu", "Am", "Cm",
    "Bk", "Cf", "Es", "Fm", "Md", "No", "Lr",
]

with MPRester() as mpr:
    docs = mpr.materials.summary.search(
        num_elements=(2, 2),
        exclude_elements=RADIOACTIVE,
        fields=["material_id"],
        num_chunks=1,
        chunk_size=1,
    )

# Yields: MPRestError 422 — "exclude_elements - String should have at most 60 characters"

What happened?

I want to query for some structures but exclude radioactive elements.

The extremely conservative MP-API settings doesn't let me exclude so many elements.

So how are people supposed to do this query?

Version

current

Which OS?

  • MacOS
  • Windows
  • Linux

Log output

(atomate2_env) ajain:~/Documents/code_venvs/atomate2_env/Part_2-MongoDB_High_Throughput$ python script2.py 
Traceback (most recent call last):
  File "/Users/ajain/Documents/code_venvs/atomate2_env/Part_2-MongoDB_High_Throughput/script2.py", line 12, in <module>
    docs = mpr.materials.summary.search(
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/ajain/Documents/code_venvs/atomate2_env/lib/python3.11/site-packages/mp_api/client/routes/materials/summary.py", line 393, in search
    return super()._search(  # type: ignore[return-value]
           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/ajain/Documents/code_venvs/atomate2_env/lib/python3.11/site-packages/mp_api/client/core/client.py", line 1456, in _search
    return self._get_all_documents(
           ^^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/ajain/Documents/code_venvs/atomate2_env/lib/python3.11/site-packages/mp_api/client/core/client.py", line 1520, in _get_all_documents
    results = self._query_resource(
              ^^^^^^^^^^^^^^^^^^^^^
  File "/Users/ajain/Documents/code_venvs/atomate2_env/lib/python3.11/site-packages/mp_api/client/core/client.py", line 795, in _query_resource
    data = self._submit_requests(
           ^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/ajain/Documents/code_venvs/atomate2_env/lib/python3.11/site-packages/mp_api/client/core/client.py", line 978, in _submit_requests
    data, total_num_docs = self._submit_request_and_process(
                           ^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^^
  File "/Users/ajain/Documents/code_venvs/atomate2_env/lib/python3.11/site-packages/mp_api/client/core/client.py", line 1237, in _submit_request_and_process
    raise MPRestError(
mp_api.client.core.exceptions.MPRestError: REST query returned with error status code 422 on URL https://api.materialsproject.org/materials/summary/?nelements_min=2&nelements_max=2&exclude_elements=Tc%2CPm%2CPo%2CAt%2CRn%2CFr%2CRa%2CAc%2CTh%2CPa%2CU%2CNp%2CPu%2CAm%2CCm%2CBk%2CCf%2CEs%2CFm%2CMd%2CNo%2CLr&_limit=1&_fields=material_id with message:
exclude_elements - String should have at most 60 characters

Activity

  1. computron commented on May 15, 2026

    @computron
    MemberAuthor

    Also, if data efficiency is desired, note that I now have to overfetch materials and then re-filter on my side. So this is the least data-efficient way to do it.

  2. tsmathis commented on May 15, 2026

    @tsmathis
    Collaborator

    The limits on the elements queries (and related queries with semi-high cardinality fields) is more of a protection thing (for the db) rather than a push for data efficiency. If unbounded inputs were allowed it would be very easy to generate heavy load on the database instance itself very quickly with these types of queries.

    That said though, I understand it's frustrating. Especially for your query which is right on the border of the current allowed character limit. So, what type of behavior/limits do you think would be reasonable here? The indexing strategy and query planning for these queries are a few years old at this point so they might be worth auditing.

  3. esoteric-ephemera commented on May 15, 2026

    @esoteric-ephemera
    Collaborator

    Just noting that filtering on the user side is very fast, this takes roughly 10s total for me (dependent on one's internet connection):

    RADIOACTIVE = {
        "Tc", "Pm", "Po", "At", "Rn", "Fr", "Ra",
        "Ac", "Th", "Pa", "U", "Np", "Pu", "Am", "Cm",
        "Bk", "Cf", "Es", "Fm", "Md", "No", "Lr",
    }
    
    with MPRester() as mpr:
        possibly_radioactive_binaries = mpr.materials.search(
            num_elements=2,
            fields=["material_id","elements"],
        )
    non_radioactive_binaries = [
        doc for doc in possibly_radioactive_binaries
        if set(ele.value for ele in doc.elements) & RADIOACTIVE == set()
    ]
  4. computron commented on May 15, 2026

    @computron
    MemberAuthor
    • On client-side filtering being fast enough: in my example, I deliberately kept the query small for reproducibility, but the same pattern at realistic scale (more compounds, more fields) would mean fetching substantially more data just to throw most of it away. Taken to its logical end, this argument applies to any server-side filter, and I think we'd agree the right place for most filtering is in the query and not in post-filtering code.
    • On the database-protection rationale: I'd push back on whether exclude_elements is actually a problem in practice. Has heavy load from this kind of query been observed, or is the limit precautionary? On an indexed field like elements, a $nin of 20-odd symbols shouldn't be significantly more expensive than the corresponding $in query, and users excluding or including a fixed chemistry set (radioactives, magnetic elements, expensive elements) is a pretty common workflow. Remember that by not providing this functionality, you are also creating database load by making the user over-fetch responses. And if the user is not doing a two-step workflow where they grab materials_ids and compositions first, filtering, and then re-fetching something more expensive for the filtered list, they will be hitting the database a lot more for data they don't actually need by doing a one step workflow where they grab all the data for everything.
    • On what limits would be reasonable: for indexed fields like elements, I'd lean toward letting users issue the query without a character cap, or at minimum raising the cap to comfortably accommodate half the periodic table's worth of two-letter symbols (~150 chars). If specific query patterns turn out to cause real load problems, those could be addressed targeted rather than via a blanket string-length check that catches legitimate use cases.
Sign up for free to join this conversation on GitHub. Already have an account? Sign in to comment

Metadata

Metadata

Assignees

No one assigned

    Labels

    bugSomething isn't working

    Type

    No type

    Projects

    No projects

      Milestone

      No milestone

      Relationships

      None yet

      Development

      No branches or pull requests

      Issue actions