Advanced Search Techniques

The free text field in the search interface does not just look for exact matches — it ranks results by relevance across several indexed fields at once. This is convenient for casual searches, but it also means the top result is not necessarily the only correct match, and a search for an exact acronym or ID can be diluted by loosely related hits.
If you need precise results, you can bypass the relevance ranking and query individual indexed fields directly. The examples below show how to do this. Just type the given strings into the search field to reproduce the described behaviour.

Searching for a specific entry

To search for a specific entry_id, enter id: followed by the value, e.g.
id:3845846
To search for a specific acronym, enter the field name entry_acronym_s: followed by the acronym:
entry_acronym_s:ECHAM3_T42_22056HMBG_ACLCOV

Wildcards

Fields whose name ends in _s (string) or _ss (string list) are stored without any tokenization, so they support the * wildcard for prefix/substring-style matching. This lets you find all entries with a similarly named value for that field, for example all versions of an acronym:
entry_acronym_s:ECHAM3_T42_22056HMBG_*
Wildcards are of limited use on the numeric (_l), date (_dt) and date-range (_rdt) fields — use the range syntax below for those instead.

Escaping special characters

Whitespace needs to be escaped with a backslash, otherwise it is interpreted as an implicit AND between two separate terms. The following searches for all versions of the CMIP5 DRS path cmip5 output1 CSIRO-QCCCE CSIRO-Mk3-6-0 1pctCO2 mon ocean Omon r1i1p1 ... tos:
entry_name_s:cmip5\ output1\ CSIRO-QCCCE\ CSIRO-Mk3-6-0\ 1pctCO2\ mon\ ocean\ Omon\ r1i1p1*tos
Whitespace is not the only character that needs care. The query syntax also treats + - && || ! ( ) { } [ ] ^ " ~ * ? : \ / as syntax, so if a value you are searching for happens to contain one of these characters (for example a colon in an identifier), escape it with a backslash as well or the query may fail to parse or match unexpectedly.

Search operators

The free text search allows the set operators AND, OR and NOT, which can be nested in parentheses:
(general_key_ss:1pctCO2 OR general_key_ss:BNU-ESM) AND (NOT id:3189360)
AND is the default operator: whitespace between two clauses behaves like AND. It is also allowed to replace NOT with a dash (-).

Ranges

Fields whose name ends in _l (the datatype long) or _dt (date) can be searched as a range. To search for datasets between 1000 and 10000 bytes:
data_size_l:[1000 TO 10000]
An open-ended bound is also allowed, for example to search for datasets larger than 1 TB:
data_size_l:[1000000000000 TO *]
You can also use an open range on both sides to check whether a field is set at all. Since projects are not "real" entries in the CERA sense, they do not have a progress acronym; the following finds all entries without one:
NOT progress_acronym_s:["" TO *]
The same works for numeric fields — note that the canonical form uses * on both ends, field:[* TO *], rather than a specific value such as 0:
NOT data_size_l:[* TO *]

Dates

Dates need to be provided in ISO format, e.g.
creation_date_dt:2017-04-06T10:39:46Z
Ranges work the same way as for numbers:
creation_date_dt:[2015-04-06T10:39:46Z TO 2017-04-06T10:39:46Z]
publication_date_dt (when the entry was published, as opposed to creation_date_dt, when the metadata record was created) can be queried the same way.

Spatial and temporal coverage

geo and date_range_rdt look like a normal field and a normal date field respectively, but they are not: querying them correctly requires specifying a relationship (does the entry's coverage contain, intersect, or lie within the given box/interval?) rather than a plain value or [A TO B] range. In the web interface, the recommended way to filter by these is the dedicated Map (bounding box) and Temporal Coverage selectors in the search sidebar, each of which lets you pick "Contains", "Intersects" or "Within" from a dropdown and takes care of the underlying syntax for you.
If you want to type the equivalent query into the free text field by hand, it looks like this:
geo:"Contains(ENVELOPE(-180,180,90,-90))"
{!field f=date_range_rdt op=Contains}[1980-01 TO 2000-01]
(op can be Contains, Intersects or Within; the coordinate order in ENVELOPE() is west, east, north, south.) This is genuinely advanced usage — get the exact syntax right, or use the sidebar widgets instead.

List of available fields

Not every field is available for every entry, and not all of them are exposed as filters in the sidebar — some are only queryable through the free text field. The suffix of a field name indicates its type:
SuffixType
_sstring
_ssstring list
_llong (number)
_dtdate
_rdtdate range
(none)special type, see the field's description
FieldDescription
idInternal numeric entry ID.
entry_acronym_sThe entry's acronym.
entry_name_sThe entry's full name/title.
entry_type_sHierarchy level as text, e.g. project, experiment, dataset_group, dataset, additional_info.
entry_type_lHierarchy level as its internal numeric code (e.g. 7 for a dataset). Mainly useful for range queries.
publication_type_sType shown in the interface, e.g. "Project", "Experiment", "Dataset".
progress_acronym_sArchival status of the entry (e.g. "completely archived"). Projects don't have one.
summary_sFree-text summary/abstract of the entry.
project_descr_ssSummary text specific to project-level entries.
format_acronym_sData format of the entry.
data_size_lSize of the entry's data in bytes.
geoSpatial coverage bounding box. See Spatial and temporal coverage.
date_range_rdtTemporal coverage. See Spatial and temporal coverage.
creation_date_dtDate the metadata record was created.
publication_date_dtDate the entry was published.
general_key_ssFree keywords assigned to the entry.
institute_info_ssInstitutes listed as a contact/owner of the entry.
pinstitute_info_ssInstitutes that are the affiliation of the entry's listed contact persons (distinct from institute_info_ss).
person_name_ssNames of persons associated with the entry.
topic_name_ssVariables/topics covered by the entry.
aggregation_descr_ssDescription of how child entries are aggregated.
code_type_ssType(s) of model/code associated with the entry.
project_name_ssName of the project the entry belongs to.
project_acronym_ssAcronym of the project the entry belongs to.
ref_type_name_sSet to "Citation-DOI" for entries that have a citable DOI.
hierarchy_ssEncodes the entry's ancestors, one string per ancestor in the form <type> @ <level> @ <acronym> @ <parent acronym>.
hierarchy_steps_ssFlat list of the acronyms of all ancestors of the entry; used to enable hierarchical filtering.
qc_institute_sOrganization associated with the entry's model/experiment metadata.
qc_experiment_sExperiment identifier associated with the entry's model metadata (e.g. a CMIP experiment name).
model_sModel name.
realm_sModelling realm (e.g. atmosphere, ocean).
variable_sVariable name.
frequency_sTemporal frequency of the data (e.g. monthly, daily).
access_sIdentifier used to build the entry's access/download link.
authors_sAuthor list.
additional_infos_ssAcronyms of any "additional info" child entries.
license_sData license.
lta_files_ssFiles stored for this entry in the long-term archive.
is_downloadableWhether at least one file of the entry can be downloaded.

Further reading

The complete, authoritative list of request parameters accepted by the search API — including parameters not exposed in the GUI, such as facet, facet.field and fl — is documented at /wdcc-api-docs.