Advanced Search Techniques
The free text field in the search interface does not just look for exact matches — it ranks results by relevance across several indexed fields at once. This is convenient for casual searches, but it also means the top result is not necessarily the only correct match, and a search for an exact acronym or ID can be diluted by loosely related hits.
If you need precise results, you can bypass the relevance ranking and query individual indexed fields directly. The examples below show how to do this. Just type the given strings into the search field to reproduce the described behaviour.
This document only covers the search box in the web interface. Developers integrating against the search API directly should refer to the interactive reference at /wdcc-api-docs, which documents the full set of supported request parameters — several of these are used internally by the map and date-range widgets described below but are not available inside the free text field itself.
Searching for a specific entry
To search for a specific
entry_id, enter id: followed by the value, e.g.id:3845846To search for a specific acronym, enter the field name
entry_acronym_s: followed by the acronym:entry_acronym_s:ECHAM3_T42_22056HMBG_ACLCOVWildcards
Fields whose name ends in
_s (string) or _ss (string list) are stored without any tokenization, so they support the * wildcard for prefix/substring-style matching. This lets you find all entries with a similarly named value for that field, for example all versions of an acronym:entry_acronym_s:ECHAM3_T42_22056HMBG_*Wildcards are of limited use on the numeric (
_l), date (_dt) and date-range (_rdt) fields — use the range syntax below for those instead.Escaping special characters
Whitespace needs to be escaped with a backslash, otherwise it is interpreted as an implicit
AND between two separate terms. The following searches for all versions of the CMIP5 DRS path cmip5 output1 CSIRO-QCCCE CSIRO-Mk3-6-0 1pctCO2 mon ocean Omon r1i1p1 ... tos:entry_name_s:cmip5\ output1\ CSIRO-QCCCE\ CSIRO-Mk3-6-0\ 1pctCO2\ mon\ ocean\ Omon\ r1i1p1*tosWhitespace is not the only character that needs care. The query syntax also treats
+ - && || ! ( ) { } [ ] ^ " ~ * ? : \ / as syntax, so if a value you are searching for happens to contain one of these characters (for example a colon in an identifier), escape it with a backslash as well or the query may fail to parse or match unexpectedly.Search operators
The free text search allows the set operators
AND, OR and NOT, which can be nested in parentheses:(general_key_ss:1pctCO2 OR general_key_ss:BNU-ESM) AND (NOT id:3189360)AND is the default operator: whitespace between two clauses behaves like AND. It is also allowed to replace NOT with a dash (-).Be careful when dropping the parentheses. Removing them and writing something like
general_key_ss:1pctCO2 OR general_key_ss:BNU-ESM -id:3189360usually behaves the same as the parenthesised version above, but mixing OR with an implicit AND or a bare NOT/- clause is a known ambiguity — the effective grouping is not guaranteed to match normal boolean-algebra precedence. When combining OR with anything else, it is safer to always group the OR clause in parentheses explicitly, as in the example above. Ranges
Fields whose name ends in
_l (the datatype long) or _dt (date) can be searched as a range. To search for datasets between 1000 and 10000 bytes:data_size_l:[1000 TO 10000]An open-ended bound is also allowed, for example to search for datasets larger than 1 TB:
data_size_l:[1000000000000 TO *]You can also use an open range on both sides to check whether a field is set at all. Since projects are not "real" entries in the CERA sense, they do not have a progress acronym; the following finds all entries without one:
NOT progress_acronym_s:["" TO *]The same works for numeric fields — note that the canonical form uses
* on both ends, field:[* TO *], rather than a specific value such as 0:NOT data_size_l:[* TO *]Dates
Dates need to be provided in ISO format, e.g.
creation_date_dt:2017-04-06T10:39:46ZRanges work the same way as for numbers:
creation_date_dt:[2015-04-06T10:39:46Z TO 2017-04-06T10:39:46Z]publication_date_dt (when the entry was published, as opposed to creation_date_dt, when the metadata record was created) can be queried the same way.Spatial and temporal coverage
geo and date_range_rdt look like a normal field and a normal date field respectively, but they are not: querying them correctly requires specifying a relationship (does the entry's coverage contain, intersect, or lie within the given box/interval?) rather than a plain value or [A TO B] range. In the web interface, the recommended way to filter by these is the dedicated Map (bounding box) and Temporal Coverage selectors in the search sidebar, each of which lets you pick "Contains", "Intersects" or "Within" from a dropdown and takes care of the underlying syntax for you.If you want to type the equivalent query into the free text field by hand, it looks like this:
geo:"Contains(ENVELOPE(-180,180,90,-90))"{!field f=date_range_rdt op=Contains}[1980-01 TO 2000-01](
op can be Contains, Intersects or Within; the coordinate order in ENVELOPE() is west, east, north, south.) This is genuinely advanced usage — get the exact syntax right, or use the sidebar widgets instead.List of available fields
Not every field is available for every entry, and not all of them are exposed as filters in the sidebar — some are only queryable through the free text field. The suffix of a field name indicates its type:
| Suffix | Type |
|---|---|
_s | string |
_ss | string list |
_l | long (number) |
_dt | date |
_rdt | date range |
| (none) | special type, see the field's description |
| Field | Description |
|---|---|
id | Internal numeric entry ID. |
entry_acronym_s | The entry's acronym. |
entry_name_s | The entry's full name/title. |
entry_type_s | Hierarchy level as text, e.g. project, experiment, dataset_group, dataset, additional_info. |
entry_type_l | Hierarchy level as its internal numeric code (e.g. 7 for a dataset). Mainly useful for range queries. |
publication_type_s | Type shown in the interface, e.g. "Project", "Experiment", "Dataset". |
progress_acronym_s | Archival status of the entry (e.g. "completely archived"). Projects don't have one. |
summary_s | Free-text summary/abstract of the entry. |
project_descr_ss | Summary text specific to project-level entries. |
format_acronym_s | Data format of the entry. |
data_size_l | Size of the entry's data in bytes. |
geo | Spatial coverage bounding box. See Spatial and temporal coverage. |
date_range_rdt | Temporal coverage. See Spatial and temporal coverage. |
creation_date_dt | Date the metadata record was created. |
publication_date_dt | Date the entry was published. |
general_key_ss | Free keywords assigned to the entry. |
institute_info_ss | Institutes listed as a contact/owner of the entry. |
pinstitute_info_ss | Institutes that are the affiliation of the entry's listed contact persons (distinct from institute_info_ss). |
person_name_ss | Names of persons associated with the entry. |
topic_name_ss | Variables/topics covered by the entry. |
aggregation_descr_ss | Description of how child entries are aggregated. |
code_type_ss | Type(s) of model/code associated with the entry. |
project_name_ss | Name of the project the entry belongs to. |
project_acronym_ss | Acronym of the project the entry belongs to. |
ref_type_name_s | Set to "Citation-DOI" for entries that have a citable DOI. |
hierarchy_ss | Encodes the entry's ancestors, one string per ancestor in the form <type> @ <level> @ <acronym> @ <parent acronym>. |
hierarchy_steps_ss | Flat list of the acronyms of all ancestors of the entry; used to enable hierarchical filtering. |
qc_institute_s | Organization associated with the entry's model/experiment metadata. |
qc_experiment_s | Experiment identifier associated with the entry's model metadata (e.g. a CMIP experiment name). |
model_s | Model name. |
realm_s | Modelling realm (e.g. atmosphere, ocean). |
variable_s | Variable name. |
frequency_s | Temporal frequency of the data (e.g. monthly, daily). |
access_s | Identifier used to build the entry's access/download link. |
authors_s | Author list. |
additional_infos_ss | Acronyms of any "additional info" child entries. |
license_s | Data license. |
lta_files_ss | Files stored for this entry in the long-term archive. |
is_downloadable | Whether at least one file of the entry can be downloaded. |
Further reading
The complete, authoritative list of request parameters accepted by the search API — including parameters not exposed in the GUI, such as
facet, facet.field and fl — is documented at /wdcc-api-docs.