4 - More SPARQL queries
This week I continue studying the basic concepts of RDF. According to my ChatGpt course, I will study other types of queries that allow to filter or aggregate results. Let’s see what that will mean.
I will again reference mainly this website for studying the concepts: https://sparql.dev/article/SPARQL_Query_Language_Basic_Concepts_and_Features.html
OPTIONAL
A user sometimes cannot be certain that the data they are looking for is present in the database. The Optional patterns allow to specify patterns that may or may not be present in the data. This is used to retrieve data that may be missing or incomplete, as it does not require the pattern to match. The syntax for an optional pattern is as follows:
OPTIONAL { query pattern }For example, if we wanted to retrieve all organisations in our dataset that have a name defined by schema:name and that may or may not have a ex:presents relation with a project , we would use the following query:
PREFIX ex: <https://example.org/kg/> PREFIX schema: <https://schema.org/> SELECT ?organisation ?organisationName ?project WHERE { ?organisation schema:name ?organisationName . OPTIONAL { ?organisation ex:presents ?project . } }can return something like:
organisation organisationName project organisation1 Organisation One project1 organisation2 Organisation Two — GROUP BY
We now look at some of the aggregation clauses when working with SPARQL. The ORDER BY clause is used to create groups of solutions, normally with the intention to perform an aggregation on each group. In the example we return the number of projects per organisation.
PREFIX ex: <https://example.org/kg/>
SELECT ?organisation (COUNT(?project) AS ?numberOfProjects)
WHERE {
?organisation ex:presents ?project .
}
GROUP BY ?organisation
This can result in something like:
| organisation | numberOfProjects |
|---|---|
| organisation1 | 4 |
| organisation2 | 2 |
| organisation3 | 7 |
- COUNT
COUNT is another aggregation functions, that allows the user to count the number of occurrences or values of the specified variable. The COUNT funtion is used in the SELECT clause.
If the user intends to count the number of projects per organisation in the database, they can use the following query:
PREFIX ex: <https://example.org/kg/>
SELECT ?organisation (COUNT(?project) AS ?numProjects)
WHERE {
?organisation ex:presents ?project .
}
GROUP BY ?organisation
- HAVING
Having is a related function to GROUP BY. The HAVING clause is used to filter the results produced by GROUP BY based on a specified condition. In the example that follows, the query returns organisations that have presented at least three projects.
PREFIX ex: <https://example.org/kg/>
SELECT ?organisation (COUNT(?project) AS ?numberOfProjects)
WHERE {
?organisation ex:presents ?project .
}
GROUP BY ?organisation
HAVING (COUNT(?project) >= 3)
- FILTER
Filters allows the user to specify additional criteria for the data they want to retrieve. A filter is used to restrict the results of the query based on certain conditions. The syntax for a filter is as follows:
FILTER(condition)
For example, if we wanted to retrieve all the people in our dataset who have a foaf:name starting with the letter “J”, we would use the following query:
PREFIX foaf: <http://xmlns.com/foaf/0.1/>
SELECT ?name
WHERE {
?person foaf:name ?name
FILTER(STRSTARTS(?name, "J"))
}
It seems that HAVING and FILTER are fairly similar.
The W3C specification describes HAVING as analogous to FILTER, but operating over groups rather than individual solutions:
- FILTER filters individual solutions/rows.
- HAVING filters groups of solutions, usually after GROUP BY and aggregation such as COUNT, SUM, or AVG.
And finally we can combine the previous concepts (GROUP BY, HAVING, COUNT and ORDER BY) together:
SELECT ?organisation (COUNT(?project) AS ?numProjects)
WHERE {
?organisation ex:presents ?project .
}
GROUP BY ?organisation
HAVING (COUNT(?project) >= 3)
ORDER BY DESC(?numProjects)