Repository navigation
[Bug]: ridiculous limits on MP-API queries #1103
Description
Activity
Also, if data efficiency is desired, note that I now have to overfetch materials and then re-filter on my side. So this is the least data-efficient way to do it.
The limits on the elements queries (and related queries with semi-high cardinality fields) is more of a protection thing (for the db) rather than a push for data efficiency. If unbounded inputs were allowed it would be very easy to generate heavy load on the database instance itself very quickly with these types of queries.
That said though, I understand it's frustrating. Especially for your query which is right on the border of the current allowed character limit. So, what type of behavior/limits do you think would be reasonable here? The indexing strategy and query planning for these queries are a few years old at this point so they might be worth auditing.
esoteric-ephemera commented
on May 15, 2026 CollaboratorMore actionsJust noting that filtering on the user side is very fast, this takes roughly 10s total for me (dependent on one's internet connection):
RADIOACTIVE = { "Tc", "Pm", "Po", "At", "Rn", "Fr", "Ra", "Ac", "Th", "Pa", "U", "Np", "Pu", "Am", "Cm", "Bk", "Cf", "Es", "Fm", "Md", "No", "Lr", } with MPRester() as mpr: possibly_radioactive_binaries = mpr.materials.search( num_elements=2, fields=["material_id","elements"], ) non_radioactive_binaries = [ doc for doc in possibly_radioactive_binaries if set(ele.value for ele in doc.elements) & RADIOACTIVE == set() ]
- On client-side filtering being fast enough: in my example, I deliberately kept the query small for reproducibility, but the same pattern at realistic scale (more compounds, more fields) would mean fetching substantially more data just to throw most of it away. Taken to its logical end, this argument applies to any server-side filter, and I think we'd agree the right place for most filtering is in the query and not in post-filtering code.
- On the database-protection rationale: I'd push back on whether exclude_elements is actually a problem in practice. Has heavy load from this kind of query been observed, or is the limit precautionary? On an indexed field like elements, a $nin of 20-odd symbols shouldn't be significantly more expensive than the corresponding $in query, and users excluding or including a fixed chemistry set (radioactives, magnetic elements, expensive elements) is a pretty common workflow. Remember that by not providing this functionality, you are also creating database load by making the user over-fetch responses. And if the user is not doing a two-step workflow where they grab materials_ids and compositions first, filtering, and then re-fetching something more expensive for the filtered list, they will be hitting the database a lot more for data they don't actually need by doing a one step workflow where they grab all the data for everything.
- On what limits would be reasonable: for indexed fields like elements, I'd lean toward letting users issue the query without a character cap, or at minimum raising the cap to comfortably accommodate half the periodic table's worth of two-letter symbols (~150 chars). If specific query patterns turn out to cause real load problems, those could be addressed targeted rather than via a blanket string-length check that catches legitimate use cases.
Code snippet
What happened?
I want to query for some structures but exclude radioactive elements.
The extremely conservative MP-API settings doesn't let me exclude so many elements.
So how are people supposed to do this query?
Version
current
Which OS?
Log output