Appendix A

Stationary Tests for Time Series

Stationarity as introduced in Sect. 5.​1, is an important property for time series analysis and forecasts, especially for applying ARIMA models (Sect.  9.​4). Since it is not obvious that a time series is stationary from the time series plots alone, statistical test are often applied to provide further supporting evidence. This section briefly discusses a specific family of stationarity tests: unit root tests.

The Augmented Dickey-Fuller (ADF) and Phillips-Perron unit root tests are two of the most popular tests for stationarity. Consider a simple ARMA model
$$\begin{aligned} \hat{L}_N= \sum _{i=1}^{p} {\psi _i}{L_{N-i}}+ \sum _{j=1}^{q} {\varphi _j}{\epsilon _{N-j}}, \end{aligned}$$
(A.1)
then the series $L_1, L_2, \ldots ,$ is stationary if the absolute value of all the roots of the polynomial
$$\begin{aligned} 1 - \sum _{i=1}^{p} {\psi _i}x^{i}, \end{aligned}$$
(A.2)
are greater than 1 (see [1]). The method tests whether the null hypothesis ‘there is a unit root’ holds. It center around the concept that a stationary series (i.e. with no unit roots) should revert to the mean and hence the lagged values should indicate relevant information for predicting the change in the series. The details are beyond the scope of this book but the test themselves are often included in the various software packages and hence are relatively simple to apply. These packages typically give significance levels at which the null hypothesis can be rejected. For further information on basics of hypothesis testing see an introduction text such as [2].
Appendix B

Weather Data for Energy Forecasting

Weather is considered one of the most important variables for forecasting energy loads. This is because many household behaviours and appliances are tied to weather, e.g.

  • When it is too cold then heating is turned on.

  • In warm countries, when it is too cold, air-conditioning is switched on.

  • How cloudy it is determines how much generation is produced from solar photovoltaics.

Combination of weather variables can change the effects. For example, temperature effects can be intensified when the humidity is high. The temperature therefore feels higher to people and thus cooling appliances may be turned on at lower actual temperatures. However, as shown in the case study in Sect. 14.​2 the effect can be unexpected. The case study showed that utilising temperature as an input did not produce more accurate low voltage level forecasts, suggesting that temperature is not a large determinant of electricity used. Possible explanations were given, in particular that most heating in the UK (the location of the data) are gas heated at this moment and therefore less likely to be influenced by temperature changes than homes which are electrically heated. However, behaviours are always changing (Sect. 13.​6.​3), and with the increased uptake of heat pumps, electrical load is likely to become more and more linked to weather effects, especially temperature.

This chapter is a short overview of some of the concepts in weather and weather forecasts because of its potentially strong links with load. Some of the main variables are discussed and the different forms they can take. For weather forecasts this section briefly discusses how they are generated so that the reader can understand potential sources of error and uncertainty.

B.1 Weather Variables and Types

Section 6.​2.​6 has already briefly introduced some of the main variables used in load forecasting. They include temperature, wind speed, humidity, solar radiance, visibility. However, other variables such as precipitation and pressure may also be useful as there are interdependencies between different weather variables which can create various effects. It should be noted that meteorology is a very complicated discipline in its own right and most data scientists will have limited knowledge of weather and climate science. For these reasons if a forecaster wishes to include more complicated weather variables, or derivations from them, it may be a good idea to consult with someone who is much more knowledgeable in this area. Having said that, the standard observed variables listed above are usually sufficient for the purposes of short term load forecasting, and additional features may only bring minimal improvements in accuracy. This section will discuss a few features to be aware of when using weather variables in your models.

One of the first things to understand is the units of your data. Although numerical weather centres typically use standard units1 for their variables there are common variations which are not always clear. The most obvious is whether the temperature is in Celsius or Fahrenheit (or maybe even Kelvin!) which may be more common depending on which country you are from. Pressure data is particularly confusing and can be written in pascals (the SI units), atmospheres, bars, PSI (pounds per square inch) or Newtons per square meter.

Each variable often has variations and some of them may be more useful than others depending on the application and context. For example, the European Centre for Medium-Range Weather forecasts (ECMWF) has a whole host of variants for each variable.2 Even an “obvious” variable like temperature has a whole host of options from which to choose from. Further, variables like radiation has variations such as longwave radiation, incident shortwave, global horizontal irradiance, direct normal irradiance etc. Although there is guidance shared by the major numerical weather prediction centres the task can still be daunting for a novice. It is usually easier to utilise observation data as it is recorded directly at a particular site. For load forecasting surface temperature and wind speed (usually split into two orthogonal components) are usually of particular interest. However, even using these there are several considerations in how to use and preprocess them (see below and Sect. B.2.2).

Although the forecast data is usually at hourly resolution (observational data may be more frequent) the form that weather variables are reported can also vary. Some values are measured instantaneously at each hour, but other variables are averaged over the entire hour. Care must be taken with the averaged values as it needs to be clear which hour interval is used to produce the final average. Is it the hour prior to the timestamped value, or the hour after, or is it formed from values from half hour either side of the value? This can have implications for how you use this variable in your model such as, for example, utilising lagged values. The instantaneous variables can also be problematic since the data may be relatively unusual at the time it is recorded compared to the rest of the hour.

It should be noted that there are many different weather products but they often come in three main forms:
  1. Observations: As the name suggests these are observed values according to various sensors. This data could be from weather stations, drifting buoys, satellites, or aircraft.

  2. Forecasts: These are developed by combining observations with a forecast model of the atmosphere. These will be described in more detail in Sect. B.2.

  3. Reanalysis: After the forecast horizon has passed, the forecast models can be re-optimised using the actual observations. This effectively gives an estimate of the past weather states.

Unlike forecasts, which are defined on a predefined grid, observations locations may be focused in particular areas and hence not near to the site of interest. This means, even if weather variables are strongly related to demand, the distance from the site may mean the observed variables are not useful for the model.

If there are several observation sites then it may be that one site improves the model more than the other, or an improvement could be produced by combining them (Sect. 13.​1). Alternatively, all the variables could be included and a method such as LASSO could be used to select the most appropriate variable inputs (Sect. 8.​2.​4). Also note that lagged values of the weather time series could be useful for the forecast models due to lagged effects (especially if the weather variables are not colocated with the demand site).

B.2 Numerical Weather Forecasts

The production of weather forecasts is known as numerical weather prediction (NWP). There are many centres around the world which create weather predictions but many are at least part funded through governments.3 NWP is an computational expensive process as it requires the optimisation of cost function requiring many variables. This section dives further into the weather prediction process and how to understand it.

B.2.1 How Weather Forecasts Are Produced

Weather prediction combines models of the atmosphere with observations, and a prior guess on the current state of the atmosphere to find a optimal estimate for the current state of the atmosphere. This current state is then evolved forward using the forecast equations to estimate the future states of the atmosphere. The process of finding the optimal state of the atmosphere is known as data assimilation, and will be described below.

The models of the atmosphere are essentially a collection of partial differential equations which includes fluid dynamics, thermodynamics and ocean-atmosphere interactions. These equations must be trained on the observed data so that they can produce forecasts for several weather variables, for the next few days (usually at hourly resolution) for every grid point in the area of interest. Many NWP centres have to generate forecast for the entire Earth. In fact the grid points are not just at latitude and longitudinal points but also go several levels above the surface. This means the forecast models must solve for a state space with an order of at least $10^{8}$. The problem is the number of observations is often an order or two smaller than this (say $10^6$). Therefore the problem is underdetermined. For this reason a prior guess is required to provide estimates for the full state space.

One main form of data assimilation is 4-dimensional variational data assimilation. It is essentially a least squares estimation (Sect. 8.​2.​1) with a regularisation term and can be written as
$$\begin{aligned} J(\textbf{x}_0) = \frac{1}{2}(\textbf{x}_0-\textbf{x}^b_0)^TB^{-1}(\textbf{x}_0-\textbf{x}^b_0) +\frac{1}{2}\sum _{k=0}^N(\mathcal {H}_k(\textbf{x}_0)-\textbf{y}_k)^T R_k^{-1} (\mathcal {H}_k(\textbf{x}_0)-\textbf{y}_k) \end{aligned}$$
(B.1)
Here the aim is to find the initial state, $\mathbf {x_0}$, that optimises the cost function above, where:
  • $\textbf{x}^b_0$ is an initial prior estimate of the initial state, also known as the background. This is usually the previous forecast value or can be a climatological estimate.

  • $\textbf{y}_k$ are the observations at time step k in the forecast horizon.

  • $\mathcal {H}_k()$ is a combination of the weather forecast model and an observation operator. This evolves the current estimate $\mathbf {x_0}$ to time step k and then transforms the state to the same location and type as the observed variable.

  • $\textbf{B}$ is the error covariance (Sect. 3.​3) for background errors.

  • $\textbf{R}_k$ is the error covariance for the observations.

Thus data assimilation is a nonlinear optimisation problem where one component measure the difference between the observations and the model forecast of the current guess and the other component measures the difference between the prior guess and the current estimate. The background term acts as a regularisation term (Sect. 8.​2.​4) and fills in the missing values from the observations. The covariance matrices act as weights for the optimisation so that the model/estimates fit closer to more precise values.

The observation operator $\mathcal {H}_k()$ is worth considering in a little more detail. Note that observations are not at grid points and usually observations are not even of the values of interest (temperature, pressure etc.). Many observations are from satellites which record information across several levels of the atmosphere. This means the observation operator must interpolate to the same location amd also transform into variables which can be compared to the state variables of interest.
Fig. B.1

Illustration of data assimilation. The final estimate (black) is found by minimising the difference between the observations (red) and the background estimate (blue). This can then be used to generate future forecasts (dotted arrow)

An illustration of the data assimilation process is shown for one variable at one grid point in Fig. B.1. The final estimate is found so that its evolved state is closest to accurate observations as well as the prior estimate (the background). The future states are generated by evolving the model even further into the future. The weather prediction models are usually assessed by considering a skill score (Sect. 7.​4).

It is worth noting that weather forecasts are not usually updated at every hour. Numerical weather forecast computations are relatively expensive and hence are usually updated once ever six hours. Hence, forecasts at one hour may come from a different model run than another hour.

Reanalysis data is essentially a forecast model equivalent of the observations. This is where the data assimilation process is retrained on the historical observations to create a best estimate of the values of the states at all the grid points. If forecast data is not available then reanalysis data can be useful as an alternative input to a load forecast model, although it should be noted that in a real world application only forecast data would be available.

B.2.2 Preprocessing

Weather data should be subject to preprocessing like most data. Since weather forecasts are defined on a grid this also invites other preprocessing approaches which may not be possible with most time series data. Below we list some ways preprocessing of weather data can be used to improve load forecasts using weather data:
  • Bias correction: As with any forecasts, weather prediction models may still retain some biases. These should be corrected before utilisation in a load forecast model to reduce the errors in the final model. The calibration of probabilistic weather forecasts can be much more complicated. This topic is investigated in detail in the book [3] which presents several methods for calibrating probabilistic forecasts.

  • Grid point selection: Since weather forecasts are defined on a grid then feature selection such as LASSO (Sect. 8.​2.​4) can be used to select which grid point produces the best forecasts.

  • Grid combination: Just as individual forecast models can be improved by combining them (Sect. 13.​1), weather forecasts could potentially be improved by combining across the nearby grid points.

  • Feature engineering within an individual time series: Exponential smoothing is a useful feature engineering technique since it can take into account the delay effects of temperature on load when used within load forecasts. For building load forecasts this creates a variable that also take into account the thermal inertia. An example of using exponential smoothing for temperature is given in [4].

  • Feature engineering with several different weather variables: Weather variables can be combined to produce new variables which may be useful for load forecasting. Two common derived values are wind chill and humidity index which can better represent the perceived cold and heat better than temperature alone. For example, wind chill takes into account the combined effect of wind speed and temperature. In cold weather a faster wind speed can make the temperature feel a lot colder than with a slower speed. This higher wind chill may potentially translate to more homes turning on heating and hence higher low voltage demands. Similarly high humidity can translate to high temperatures feeling hotter than when humidity is lower.

  • Feature engineering across the grid: The grid of data points around a location contain much more information than simply taking an average. The variation across the grid provides further information as well as describes some of the dynamics within the area. Browell and Fasiolo use a feature extraction technique to produce probabilistic net load forecasts in [5]. They derived features such as max, min and standard deviations from the gridded data as inputs to their load forecast models.

Appendix C

Load Forecasting: Guided Walk-Through

This section will consider a forecasting trial and how to select inputs to the models, test some of the different forecasts presented in this book, as well as evaluate their accuracy. The context of this walk-through will be to generate day ahead forecasts (with the forecast origin starting at the beginning of the day). In an ideal situation the forecast models will be retrained at the start of each new day on the extra available data, but for this task only train the data once on the training data (for use on the validation set) and once again on the combined training and validation set (for use on the test set), to reduce the computational cost of constant retraining. Consider other training approaches in future experiments. The steps here will largely follow those presented in Chap. 12.

To run through the assessment select an open dataset. There is a list of data available from here https://​low-voltage-loadforecasting.​github.​io/​ but the following are also possible choices:
  • Irish Smart Meter data, available from https://​www.​ucd.​ie/​issda/​data/​commission-forenergyregulat​ioncer/​. This is one of the most commonly used smart meter data sets.

  • London Smart Meter data consisting of half hourly demand data from over 5,500 smart meters in London and is readily downloadable from the kaggle website https://​www.​kaggle.​com/​jeanmidev/​smart-meters-in-london. It also consists of London weather data for several variables including temperature

  • The Global Energy Forecasting Competition (GEFCOM) 2014 data [6]. This data set is at a higher voltage level than smart meter data but does consist of several years of demand including temperature data. Hence although it is smoother and more regular than LV level demand, there is a lot of data for training and testing.

To recreate the analysis in this book, individual smart meter data can be aggregated to simulate LV feeder level data. This will obviously be smoother than the smart meter data but may make the demand easier to predict. Once the data has been collected, first start by exploring patterns and relationships in the data. This is often one of the most important parts of the model development.

  1. Quick checks: Before doing anything perform a quick check that there isn’t too many missing values in the data (Sect. 6.​1). If the data has too many gaps then cut the data down to a shorter dataset with less than 5% missing values. If using the smart meter data, then choose a few hundred which have at least 95% of their values. The GEFCOM data should be relatively clean.

  2. Split the data: As shown in Sect. 8.​1.​3 the data needs to be partitioned into training, validation and test sets. Split the time series into the oldest $60\%$ set of data as the training data, the next $20\%$ as the validation and the final $20\%$ as the test set. Other split ratios can be tested but this is sufficient for this initial study.

  3. Plot the time series: Plot the training data time series (Sect. 6.​2.​2), or in the case of the smart meter data, plot several of them to get a better understand of the structures in the data. Notice any patterns, is there annual seasonality? If there is, when is the demand highest and lowest? What about any trends, is the demand increasing or decreasing over time? Zoom in on a few weeks of data, is their daily or weekly seasonality? Are the differences obvious for different days of the week? If there is no obvious seasonality then does it look stationary? In this case, you may want to consider applying a unit root test (see Appendix A).

  4. Seasonal correlations: For the data with suspected seasonality, plot the autocorrelation and partial autocorrelation plots (see Sect. 3.​5). If considering a lot of smart meter data then it won’t be practical to check each ACF and PACF plot so instead consider ways of aggregating the information. As was given in the Case Study (Fig. 14.​2, Sect. 14.​2), you could draw a scatter plot of the ACF and PACF for different lags on the x and y axis which may be significant, e.g. the daily and weekly lags (48 and 336 if considering half-hourly data). This will show a range of different seasonalities. Focus on particular smart meters which have the strongest and weakest correlations and plot their ACF and PACF and compare this to their actual smart meter time series. It may reveal unusual behaviour not previously expected. Alternatively, consider the autocorrelation and partial autocorrelation of the average smart meter profile. When considering the ACF and PACF where are the lags strongest? Is the correlation at lag 336 (a week) larger than the daily lags (48)?

  5. Idenifying anomalous values: In addition to the missing values (point 1), what are some other anomalous values you can find (Sect. 6.​1)? These may have been visible from the time series plot (point 3), showing very large values, or negative values (which perhaps should not be in demand data—unless there is some solar generation connected to the household and the smart meters are showing net demand). You may want to create a simple seasonal model to identify outliers. For example, if there is annual seasonality then for half hourly data, fit a simple seasonal model of the form
    $$\begin{aligned} a+b + \sum _{p=1}^P c_p \sin \left( \frac{2\pi p t }{24 \times 365} \right) + d_p \cos \left( \frac{2\pi p t }{24 \times 365} \right) , \end{aligned}$$
    (C.1)
    where $P=2$ or 3. Taking the difference between the model and the data should remove the annual seasonality. If there is a linear trend in the series then update (C.1) with additional linear terms. With the resultant residual series, check which points are more than three standard deviations from the mean and select these to be replaced. Of course a more sophisticated model could be chosen using other features discovered in the analysis (for example by including weekly and daily components like the ST model in Sect. 14.​2.​3) but for now just consider this approach.
  6. Pre-processing: For the missing and anomalous values identified in point 1 and 5 they should be replaced with appropriate values. For the data with daily/weekly seasonality impute (Sect. 6.​1.​2) the anomalous and missing values in the training and validation set with averages of the weekly values at the same time period before and after this value. In other words for a missing value at 2 PM on a Tuesday, take the 2 PM on the Tuesday before and the 2 PM the Tuesday after. If there are bigger gaps so that these values are also missing, several runs of the imputation may be required. Since the data chosen is relatively clean this should be sufficient.

The above procedure should produce a relatively clean data set with no missing values which will makes extra analysis much simpler and facilitate creating the forecast models. In addition, the basic analysis above will have highlighted some of the core features of the data, in particular whether the data is stationary and different types of seasonality. The next step is to start a more detailed analysis of other features of the data and walk through choosing a few models to train and test on the hold-out test set. These are described by the following points.

  • Visualisation of explanatory variables: Start to explore the other relationships in the training set data. The autocorrelations and large patterns and trends in the data have already been considered. If there are other explanatory variables available with the main demand data, such as temperature, then consider some scatter plots, or if there is lots of potential explanatory variables then consider a pair plot (see Sect. 6.​2.​2). What do the relationships look like? Are they linear, or nonlinear. If they look nonlinear would a simple quadratic or cubic relationship describe it accurately? Do different hours of the day have different levels of seasonality, or stronger correlations with the explanatory variables? Try plotting the cross correlation (Sect. 6.​2.​3), when do the largest values occur? This will show if some of the lagged values are also important to include in the models. Also consider the difference in the days of the week. Plot an average profile for the different days and see if there is similarities or differences which indicate they should potentially be treated specifically within the models, e.g. through dummy variables (as seen with multiple linear regression in Sects. 9.​3 and 14.​2.​3). When modelling potential relationships consider the adjusted $R^2$ (Sect. 6.​2.​3), which variables give the largest value?

  • Initial model selection: Choose a selection of models which may be suitable based on the data analysis. If the data has strong autocorrelations and linear relationships with explanatory variables then perhaps generate a linear model or ARIMAX type model (Sect. 9.​4). Consider a number of different choices for the explanatory variables: perhaps linear, quadratic and cubic versions. If there are strong seasonalities in the data, then consider the seasonal exponential smoothing models (Sect. 9.​2). If there is strong explanatory variables but no seasonal relationship then consider a deep neural networks as in Sect. 10.​5 or support vector regression (Sect. 10.​2).

  • Benchmark models: Pick some simple benchmark models based on the features observed in the analysis (see Sect. 9.​1 for common benchmarks). If there is seasonalities in the data then consider a seasonal persistence model, or take a seasonal moving average model. If there isn’t any seasonality then consider just a simple persistence model. Some of the simpler models chosen in the previous step can also be considered benchmarks. Alternatively base some of the benchmarks on a single explanatory variable. The importance of the benchmarks is to understand some of the main features which are important to the forecast accuracy and suggest potential improvements. Hence simpler benchmarks can be more informative them more sophisticated comparisons. On the other hand comparison to the state-of-the-art can be an extremely useful litmus test for the quality of your forecast.

  • Error measures: Choose at least one error measure which will best represent the accuracy of your forecast (Chap. 7). If dealing with smart meter data, or data which has many small values then MAPE is unsuitable as the errors will be inflated for smaller values (or not defined!). If you wish to compare the accuracy across several time series then choose relative error measures (such as normalised versions of RMSE or MAE, or even MAPE as long as the values are not small).

  • Training: Train the selected models (and benchmarks) on the observations in the training data (Sect. 8.​2), this will also mean training a range of the same type of models (Neural Networks, multiple linear regressions, etc.) with different parameters (for example for neural networks this will mean trying different numbers of nodes and layers, for multiple linear regression using different input variables, possibly including transformations of those variables). There are ways of automatically selecting the inputs or reducing possible overtraining, notably by using regularisation (Sect. 8.​2.​4) and in the case of likelihood based models, using information criteria (Sect. 8.​2.​2). However, in this walk-through just consider cross-validation (Sect. 8.​1.​3) for choosing the hyper-parameters and selecting the final models to test. Most standard packages will train the models according to standard measures/loss functions (for example, multiple linear regression will use least squares—(Sect. 8.​2)) however perhaps change the target measure to better fit the error measure chosen in the last part.

  • Model/Hyper-parameter selection: Use the trained models to forecast on the validation set (Sect. 8.​1.​3) and compare the errors. Select a couple of models from each family which perform the best (have the smallest errors). Keep all the benchmarks, but use this opportunity to see which models appear to have the smallest forecast errors. Which ones are more accurate than the benchmarks, and by how much? This will be interesting to see if there is any change when comparing with the test set. There is usually many models, and variants of the same model, tested on the validation set and therefore there is a possibility a model is the most accurate simply by chance alone.

  • Forecasting: Retrain the data on the combined training and validation set. Now produce the day ahead forecasts over the test set!

  • Evaluate the results: Calculate the errors for each method. Now is the time to start to evaluate the results and better understand what are the differences, the similarities, and what are the core features which make one model better than another. Of the most accurate methods what are the common features, are any of these features in the least accurate methods? What features are missing in the worst performing methods? Does the inclusion of particular explanatory variables improve a method? Do different models have different accuracies for different times of the day?

You’ve now completed a full forecasting trial! However there are several ways to improve the accuracy and quality of your models. Here are some further things to try:
  • Improvements based on the error analysis: Based on your evaluation of the results is there obvious ways to improve the results? Do any explanatory variables show any improvement when included versus not included? If they are not included in the most accurate model then see if they improve them when added. Is there a feature of the best model that is not in the other models? Perhaps if this feature is included in other models they will have greater accuracy compared to the current best model?

  • Residual analysis: Plot the residual time series, is there any prominent features in the series, seasonalities or trend? If so then update the model to include this. The residual series should be stationary, check with a unit test (Appendix A). Is there any correlation remaining in the residuals of the most accurate models. If so then, as shown in Sect. 7.​5, there is a simple way to improve a forecast by adding further autoregressive components to the models.

  • Combining models: A common way to improve any individual forecast model is to combine several forecast models together as shown in Sect. 13.​1. Generate a new forecast by taking a simple average over some (or all) of the models used in the test set.

  • Feature extraction: In the above study features were only chosen by comparing different models in the validation set. Now consider more automated methods for selecting the features. Select an ARIMA model, using the Akaike Information Criteria (Sect. 8.​2.​2). Also consider a Linear Model which uses all available features (and derivations of them) and choose the final variables via a LASSO model (Sect. 8.​2.​4). Also consider creating models using other methods but with the selected variables from the LASSO.

  • Probabilistic models: If considering a dataset with lots of historical data, then probabilistic forecasts are valid ways to improve the estimates of the future demand. A host of methods were introduced in Chap. 11. Take the linear model used for the point forecast and use it within a quantile regression for $0.05, 0.1, \ldots , 0.95$ quantiles as described in Sect. 11.​4. Residual bootstrapped forecasts are also relatively simple to implement for 1-step ahead forecasts (Sect. 11.​6.​1). Using the residual errors add adjustments to each forecast at each time step to create a new realisation of the future demand. Repeat this several hundreds (or preferably thousands) of times to produce a ensemble of forecasts. Now generate empirical quantiles (Sect. 3.​4) at each time step of the day to generate another quantile forecast (Sect. 5.​2). Probabilistic forecasts require a probabilistic forecast measure such as the Pinball loss score and the continuous ranked probability score (CRPS) as introduced in Chap. 7.

  • Your own investigation: Outside of this book there is a wealth of models and methods which haven’t been covered. To get you started more models and methods can be found from the further reading in Appendix D.1 and D.2.

Appendix D

Further Reading

This book has covered a large number of topics and is an introductory text to forecasting for low voltage electrical demand series. It should provide the reader with sufficient information to design your own trials and implement your own forecasts methods. Whatever your level of knowledge hopefully there is enough methods and techniques to teach you something new, but of course there is so much more research and insightful material out there. This section outlines some further reading which may be useful for expanding on many of the topics in this book as well as other methods which were not discussed.

Note that in some cases, especially concerning energy data, the sources are websites which may be subject to change.

D.1 Time Series Analysis and Tools

Chapters 5–7 covered a wide range of techniques for measuring forecast errors, analysing relationships and extracting features from the data, and methods for model selection. As a general resource, https://​robjhyndman.​com/​hyndsight/​ is a highly recommended blog by one of the leading experts in time series forecasting, Prof Rob J Hyndman, which features lots of information on forecasting methodologies and some of the latest research for time series. In addition he has published a free e-book [7]4 which presents further details on times series forecasting principles and techniques, as well as examples of their implementations in R. There are plenty of books out there on time series analysis but the authors have found the book by Ruppert and Matteson [1] to be an excellent resource.

For the interested reader the following is some further reading on some of the specific topics covered in Chaps. 5–7.

  • Detection of outliers is a vast topic but a sophisticated algorithm for detecting outliers can be found [8] which has an accompanying R package called stray.5

  • Probabilistic scoring functions is a rapidly developing field. A detailed and advanced look at these can be found in [9]. A more accessible introduction on the concepts of calibration and sharpness can be found in [10]. A interesting study comparing various probabilistic scoring functions for multivariate data can be found in [11]. An example of using Energy Scores for ensemble forecast evaluation is given in [12] for an offshore wind forecasting application.

  • Bias-variance trade-off is one of the most important topics in machine learning and forecasting. A popular book on machine learning with a very readable and accessible introduction is [13] which has been made freely available online.6 This book also has a good overview of many other topics in data science.

D.2 Methods: Time Series and Load Forecasting

Chapters 9–11 has only given a broad overview of forecasting models for time series, especially with regard to probabilistic methods. For the interested readers there is extensive resources available for learning more about time series forecasting, in particular for energy systems. This section will outline some of them. Further reading for forecasting specifically for LV systems will be given in Appendix D.3.

In addition to Rob Hyndman’s blog as mentioned in the Appendix D.1, some of the latest research in energy forecasting can be found on Tao Hong’s blog, http://​blog.​drhongtao.​com/​, in which he posts regularly about topics in energy forecasting as well as about the Global Energy Forecasting Competition (GEFCom)7 which he co-organises. Tao Hong is also the co-author on a useful introduction to Probabilistic load forecasting [14]. Reference [15] is a great text for implementing machine learning models in python, including accessible introductions to random forest, support vector machines, artificial neural networks, boosting and reinforcement learning. Although not specifically focused on energy, the large open access compendium by Petropoulos et al. [16] provides descriptions of a whole array of topics, and further links to additional literature.

Probabilistic forecast are becoming more common and as a result more packages are becoming available which are specialised to this. For example, ProbCast is an R package which provides a number of probabilistic forecasting methods as well as visualisation and evaluation functionality [17]. This includes implementations of parametric and nonparametric methods and Gaussian copulas.

In the following references there are further details, examples and theory on some of the forecast methods and techniques covered in this book.

  • Exponential smoothing: An excellent chapter on exponential smoothing can be found in [7]. Holt-Winters-Taylor double seasonal exponential smoothing was first introduced in [18] where it was also applied to short term electricity demand forecasting.

  • ARIMAX methods: Reference [7] also has an overview of ARIMAX models including the seasonal variants, SARIMA. A detailed investigation into ARIMAX models can also be found in [1].

  • ANN: Reference [13] provides an detailed introduction to neural networks, including the technique of back-propagation for training them.

  • Classic machine learning methods: An in-depth overview of classic approaches (though note, not focused on time series forecasting specifically) like additive models, boosting and random forests, support vector machines, and nearest neighbour methods, see the book by Hastie, Tibshirani and Friedman [19].

  • Random forest: Readers may be interested in reading papers by one of the originators of Random Forests [20]. Examples of them used in short term load forecasting can be found in [21].

  • Support vector regression: Reference [22] serves as a good introduction and overview to support vector regression.

  • LASSO: Reference [23] is an excellent example of appling LASSO models to energy forecasting in the GEFCom2014 competition.

  • Generalised Additive Models: Generalised additive models were the major components of the two winning models in the Probabilistic load forecasting track of the GEFCom2014 competition. The winning paper [24] is definitely worth a look and considers a quantile regression form for the forecast. For the reader interested into diving deeper in GAMs, the introductory book by Simon N. Wood is invaluable [25] who also developed the mgcv (Mixed GAM Computation Vehicle with Automatic Smoothness Estimation) package in R. We also recommend the free online resource by Christoph Molnar on “Interpretable Machine Learning” [26] which has an excellent breakdown of GLMs and GAMs and their pros and cons.

  • k-nearest neighbours: An application for short term load forecasting is given in [27] More specific to the topic of this book this paper gives an example for low voltage demand forecasting [28, 29].

  • Gradient-boosted regression trees: The tutorial in [30] gives an overview of the basic Gradient Boosting Machine and its most important hyper-parameters. See the documentation of XGBoost8 and LightGBM9 for the documentation of the most popular implementations, their interfaces for different programming languages and their respective hyperparameters.

  • Quantile regression and kernel density estimation: An example of each of these methods including an example of how to combine them to produce an improved forecast is given in [31] for medium term probabilistic load forecasts applied in the GEFCom2014 competition.

  • Copula’s and GARCH models are common in financial applications hence [1] provides an excellent and accessible overview into both techniques and gives many examples of implementations of the methods in R. A comprehensive look at copula’s is given in [32].

  • Deep learning approaches: A more in-depth overview of deep learning models (not focused on time series), is available in the free book by Goodfellow, Bendio and Courville [33]. For more specialised deep learning approaches to time series forecast see the papers on DeepAR [34] and N-BEATS [35] as well as [36, 37].

There was additional techniques and topics discussed in Chap. 13, including model combination and hierarchical forecasting. Below is some additional reading on these topics:
  • Combining Forecasts: This topic was introduced in Sect. 13.​1. Armstrong [38] gives several insights into combining forecasts. An advanced method using copula’s to combine forecasts is outlined in [39]. For the application of load forecasting an example of combining probabilistic load forecasts is given in [40]. Rob Hyndman has recently written a review of forecast combination over the last 50 years for those interested in the variety of techniques which are available [41].

  • Statistical Significance Tests. The Diebold-Mariano test was introduced in Sect. 13.​5. A detailed example of using the Diebold-Mariano test for probabilistic forecasts is given in [11]. The paper by Harvey, Leybourne, and Whitehouse [42] dives into significance tests in more detail, including some of the drawbacks and alternative methods available.

  • Hierarchical Forecasting: An area which has gained increasing interest is hierarchical forecasting. As very briefly introduced in Sect. 13.​2, this involves forecasting at different levels whilst considering coherence between them. This topic is considered in [43] for smart meter data. The GEFCom201710 included a hierarchical component where the aim was to produce forecasts of zones and then the total load of ISO New England [44]. A brilliant introduction is also available in [7].

  • Special Days: As discussed in Sect. 13.​6.​2 special days can have very different behaviour than expected. The authors in [45] present methodologies for forecasting the load on Special days in France.

  • Calibrating Probabilistic Forecasts: Processing of probabilistic forecasts was only briefly mentioned in Sect. 7.​5. These techniques are also important for preprocessing weather forecasts before they are used in load forecasting (Sect. B.2.2). The book [3] presents several methods for calibrating probabilistic forecasts.

  • Other Pitfalls: The authors of Hewamalagea et al. [46] present common pitfalls and best practice with time series forecast evaluation. This includes benchmarks, error measures, statistical significance, and issues caused by trends, heteroscedasticity, concept shift/drift, outliers and many others.

D.3 Low Voltage Forecasting Examples

The case study introduced in Chap. 14 has highlighted a real forecast example for low voltage applications. There is numerous other examples of forecasting at the low voltage level, most of it for individual smart meter (household) level demand. Smart meter rollouts means there are many more forecast papers examining household level demand and some of these are listed below. The following review paper by two of the authors of this book investigate the methods, explanatory variables, and applications in LV forecasting [47]. This will help any readers get up to speed with the vast amount of techniques and methods being applied in this area, many of which are described in this book.

The forecast case study Chap. 14 is largely based on a piece of research by one of the authors and the interested readers is encouraged to read the original papers for further details and extra analysis. The low voltage feeder forecast work is based on the paper [48] that also includes the application of various kernel density estimation methods, which were not presented in the Chap. 14.

The literature for low voltage level demand is quite sparse and is mainly focused on aggregations of smart meter data. An interesting implementation of hierarchical probabilistic forecasts is presented in [49, 50] which considers several levels of aggregations of the smart meters. The methods are applied whilst also making the forecasts coherent with each other (ensuring the sum of the forecasts equal the forecast of the aggregation). The example of support vector regression (Sect. 10.​2) and random forest regression (Sect. 10.​3.​2) shared in the storage control in Sect. 15 is based on the paper in [51]. An example of k-nearest neighbours for forecasting low voltage demand is given in [28].

The majority of smart meter forecasting has been for point forecasts but since household demand is quite volatile it can be quite difficult to model accurately. An interesting example of additive models applied to creating probabilistic household forecasts is given in [52], which also includes applications to household batteries. In [53] the authors use partially linear additive models (PLAMs) with an extension to ensure that volatile components of the demand can be modelled. PLAMs are extensions to generalised additive models (Sect. 9.​6). Since GAMs are usually restricted to smooth components they aren’t necessarily appropriate for household level demand, and is why extensions may be appropriate and necessary. The authors in [54] apply a convolutional neural network (Sect. 10.​5) approach to forecasting residential demand.

Due to the volatility of smart meter data, probabilistic forecasting is becoming increasingly common. An excellent example of short term probabilistic smart meter load forecasting for kernel density estimation and probabilistic Holt-Taylor-Winters forecasts is given in [55]. The authors in [56] present an example of producing probabilistic forecasts for smart meters via a form of quantile regression (Sect. 11.​4) using a gradient boosted method.

Since peaks are often one of the main interesting features in low voltage demand, an important area of focus is on peak demand forecasting which falls naturally within the sphere of Extreme Value Theory. An excellent introduction to these methods and their application to low voltage level demand are given in the book by Jacob et al. [57].

Finally for something a little more esoteric the reader is pointed in the direction of [29, 58, 59]. As shown briefly in Sect. 13.​3, due to the volatile and spiky nature of smart meter data, peaks may shift in position (someone getting home later from work or university will shift their behaviour accordingly). This means that traditional pointwise methods for measuring errors (like MAPE, MAE and RMSE introduced in Chap. 7 of this book) may not be appropriate due to the ‘double penalty effect’. For a peak which is missed slightly (e.g. half an hour too early) will be penalised twice: once for missing the real peak and secondly for the forecast of a peak which didn’t occur. The above papers dive further into the adjusted error measure given in Sect. 13.​3 and also consider other updates and variants.

D.4 Data and Competitions

The best way to develop a deep understanding about forecasting and applying various techniques is to start coding up and practicing with real data. This section will outline a few publicly available datasets.

One way data is made available is through competitions. In recent years websites such as Kaggle (https://​www.​kaggle.​com/​) have hosted competitions for machine learning based problems. Competitions provide opportunities to compete against other participants, trial new methods and learn more about what makes a good forecast. The paper [60] gives a brilliant overview of their history and some of the major learnings from forecasting competitions. Some of the major time series forecasting competitions are
  • The M-Competitions, one of the first time series forecasting competitions, starting with M1 in 1982. Each competition often has increasing numbers of competitors, complexity and the number of time series. The M4 competition in 2018 consisted of 100, 00 time series to forecast [61].

  • The Global Energy Forecasting Competition. A forecasting competition started in 2012 focused on energy demand forecasting, the second in 2014 focused on probabilistic energy forecasting [6]. The 2017 competition focused on hierarchical aspects of energy forecasting.

The Global Energy Forecasting Competition review papers are a brilliant resource for learning about some of the state-of-the-art methods in energy forecasting [6, 62]. See also specific papers on some of the methods, particularly in the probabilistic tracks [24, 31]. The data has also been made available for the reader to try their own forecasting methodology (see Tao Hongs blog for this data and others: http://​blog.​drhongtao.​com/​2016/​07/​datasets-for-energy-forecasting.​html).

Unfortunately the GEFCom data is typically of a higher voltage than the topic discussed in this book. Hence, although they are good for developing your first load forecasting methods, the demand is typically much smoother than that which is of interest in this book. Low voltage data is actually much sparser and due to this, publicly available data sets have often been overanalysed, increasing the possibility of them being subject to biases and unrealistic foreknowledge about the dataset.

The authors have compiled a list of LV data sets which may be useful for the readers own investigations.11 However, below is a specific set of publicly available data contained on that list which may be of use for designing your own models:
  1. Irish Smart Meter Data: This is one of the first publicly available sets of smart meter data. It consists of data from about 4000 smart meters. https://​www.​ucd.​ie/​issda/​data/​commissionforene​rgyregulationcer​/​.

  2. London Smart meter data. Contains half hourly demand data for over 5500 smart meters in London as well as local weather data. https://​www.​kaggle.​com/​jeanmidev/​smart-meters-in-london.

  3. REFIT: High resolution (8 s) electrical load data set for 20 households including appliances. https://​pureportal.​strath.​ac.​uk/​en/​datasets/​refit-electrical-load-mea-surements-cleaned.

  4. Open Power System Data Platform: Has data on households data at various resolutions including solar data. Also an important resource for other power system data such as price, weather and demand. https://​data.​open-power-system-data.​org/​.

  5. UK-DALE: high resolution (6 s) data for five households including individual appliance data. https://​jack-kelly.​com/​data/​.

  6. GREEND: 1 s resolution data from 8 households. https://​sourceforge.​net/​projects/​greend/​.

  7. Behavioural Energy Efficiency—15min resolution data for 200 households. https://​zenodo.​org/​record/​3855575.

Notice that the majority of this data is smart meter/household level since unfortunately there is very little low voltage data. However a estimate of a LV network can be simulated via aggregations of smart meters although it should be noted that these are slightly different and therefore the equivalence isn’t exact [48].

References
  1. 1.
    D. Ruppert, D.S. Matteson, Statistics and Data Analysis for Financial Engineering: With R Examples. Springer Texts in Statistics (2015)
  2. 2.
    F.M. Dekking, C. Kraaikamp, H.P. Lopuhaä, L.E. Meester, A Modern Introduction to Probability and Statistics: Understanding Why and How (Springer, London, 2005)
  3. 3.
    S. Vannitsem, D.S. Wilks, J.W. Messner (eds.) Statistical Postprocessing of Ensemble Forecasts (Elsevier, 2018)
  4. 4.
    T-H. Dang-Ha, F.M. Bianchi, R. Olsson, Local short term electricity load forecasting: automatic approaches, in 2017 International Joint Conference on Neural Networks (IJCNN) (2017), pp. 4267–4274
  5. 5.
    J. Browell, M. Fasiolo, Probabilistic forecasting of regional net-load with conditional extremes and gridded NWP. IEEE Trans. Smart Grid 12(6), 5011–5019 (2021)Crossref
  6. 6.
    T. Hong, P. Pinson, S. Fan, H. Zareipour, A. Troccoli, R.J. Hyndman, Probabilistic energy forecasting: global energy forecasting competition 2014 and beyond. Int. J. Forecast. 32, 896–913 (2016)Crossref
  7. 7.
    R.J. Hyndman, G. Athanasopoulos, Forecasting: Principles and Practice, 2nd edn. (OTexts, Melbourne, Australia, 2018). https://​Otexts.​com/​fpp2. Accessed on July 2020
  8. 8.
    P.D. Talagala, R.J. Hyndman, K. Smith-Miles, Anomaly detection in high-dimensional data (2019)
  9. 9.
    T. Gneiting, A.E. Raftery, Strictly proper scoring rules, prediction, and estimation. J. Am. Stat. Assoc. 102, 359–378 (2007)
  10. 10.
    T. Gneiting, F. Balabdaoui, A.E. Raftery, Probabilistic forecasts, calibration and sharpness. J. Roy. Stat. Soc.: Ser. B (Stat. Methodol.) 69(2), 243–268 (2007)MathSciNetCrossrefzbMATH
  11. 11.
    F. Ziel, K. Berk, Multivariate forecasting evaluation: on sensitive and strictly proper scoring rules (2019)
  12. 12.
    C. Gilbert, J. Browell, D. McMillan, Probabilistic access forecasting for improved offshore operations. Int. J. Forecast. 37(1), 134–150 (2021)Crossref
  13. 13.
    C.M. Bishop, Pattern Recognition and Machine Learning (Information Science and Statistics) (Springer, Berlin, Heidelberg, 2006)
  14. 14.
    T. Hong, S. Fan, Probabilistic electric load forecasting: a tutorial review. Int. J. Forecast. 32(3), 914–938 (2016)Crossref
  15. 15.
    A. Gron, Hands-On Machine Learning with Scikit-Learn and Tensor Flow: Concepts, Tools, and Techniques to Build Intelligent Systems, 1st edn. (O’Reilly Media, Inc., 2017)
  16. 16.
    F. Petropoulos, D. Apiletti, V. Assimakopoulos, M.Z. Babai, D.K. Barrow, S. Ben Taieb, C. Bergmeir, R.J. Bessa, J. Bijak, J.E. Boylan, J. Browell, C. Carnevale, J.L. Castle, P. Cirillo, M.P. Clements, C. Cordeiro, F. Luiz Cyrino Oliveira, S. De Baets, A. Dokumentov, J. Ellison, P. Fiszeder, P.H. Franses, D.T. Frazier, M. Gilliland, M.S. Gönül, P. Goodwin, L. Grossi, Y. Grushka-Cockayne, M. Guidolin, M. Guidolin, U. Gunter, X. Guo, R. Guseo, N. Harvey, D.F. Hendry, R. Hollyman, T. Januschowski, J. Jeon, V.R.R. Jose, Y. Kang, A.B. Koehler, S. Kolassa, N. Kourentzes, S. Leva, F. Li, K. Litsiou, S. Makridakis, G.M. Martin, A.B. Martinez, S. Meeran, T. Modis, K. Nikolopoulos, D. Önkal, A. Paccagnini, A. Panagiotelis, I. Panapakidis, J.M. Pavía, M. Pedio, D.J. Pedregal, P. Pinson, P. Ramos, D.E. Rapach, J.J. Reade, B. Rostami-Tabar, M. Rubaszek, G. Sermpinis, H.L. Shang, E. Spiliotis, A.A. Syntetos, P.D. Talagala, T.S. Talagala, L. Tashman, D. Thomakos, T. Thorarinsdottir, E. Todini, J.R. Trapero Arenas, X. Wang, R.L. Winkler, A. Yusupova, F. Ziel, Forecasting: theory and practice. Int. J. Forecasting 38(3), 705–871 (2022)
  17. 17.
    J. Browell, C. Gilbert, Probcast: open-source production, evaluation and visualisation of probabilistic forecasts. 5 (2020)
  18. 18.
    J.W. Taylor, Short-term electricity demand forecasting using double seasonal exponential smoothing. J. Oper. Res. Soc. 54, 799–805 (2003)CrossrefzbMATH
  19. 19.
    T. Hastie, R. Tibshirani, J. Friedman Data Mining, Inference, and Prediction. The Elements of Statistical Learning. Springer Series in Statistics (2009)
  20. 20.
    L. Breiman, Random forests. Mach. Learn. 45(1), 5–32 (2001)CrossrefzbMATH
  21. 21.
    G. Dudek, Intelligent Systems’2014. Advances in Intelligent Systems and Computing. Short-Term Load Forecasting Using Random Forests, vol. 323 (Springer, Cham, 2015), , pp. 821–828
  22. 22.
    A.J. Smola, B. Schölkopf, A tutorial on support vector regression. Stat. Comput. 14, 199–222 (2004)MathSciNetCrossref
  23. 23.
    F. Ziel, B. Liu, Lasso estimation for gefcom2014 probabilistic electric load forecasting. Int. J. Forecast. 32(3), 1029–1037 (2016)Crossref
  24. 24.
    P. Gaillard, Y. Goude, R. Nedellec, Additive models and robust aggregation for gefcom2014 probabilistic electric load and electricity price forecasting. Int. J. Forecast. 32(3), 1038–1050 (2016)Crossref
  25. 25.
    S.N. Wood, Generalized Additive Models: An Introduction with R, 2nd edn. (Chapman and Hall/CRC, 2017)
  26. 26.
    C. Molnar, Interpretable Machine Learning: A Guide For Making Black Box Models Explainable (Independently published, 2022)
  27. 27.
    A.T. Lora, J.M. Riquelme Santos, J. Cristóbal Riquelme, A. Gómez Expósito, J. Luís Martínez Ramos, Time-series prediction: application to the short-term electric energy demand, in Current Topics in Artificial Intelligence ed. by R. Conejo, M. Urretavizcaya, J-L. Pérez-de-la Cruz (Springer, Berlin, Heidelberg, 2004) pp. 577–586
  28. 28.
    O. Valgaev, F. Kupzog, H. Schmeck, Low-voltage power demand forecasting using k-nearest neighbors approach, in 2016 IEEE Innovative Smart Grid Technologies - Asia (ISGT-Asia) (2016), pp. 1019–1024
  29. 29.
    M. Voß, A. Haja, S. Albayrak, Adjusted feature-aware k-nearest neighbors: utilizing local permutation-based error for short-term residential building load forecasting, in 2018 IEEE International Conference on Communications, Control, and Computing Technologies for Smart Grids (SmartGridComm) (IEEE, 2018), pp. 1–6
  30. 30.
    A. Natekin, A. Knoll, Gradient boosting machines, a tutorial. Front. Neurorobot. 7, 21 (2013)Crossref
  31. 31.
    S. Haben, G. Giasemidis, A hybrid model of kernel density estimation and quantile regression for gefcom2014 probabilistic load forecasting. Int. J. Forecast. 32, 1017–1022 (2016)Crossref
  32. 32.
    D. Kurowicka, R. Cooke, High-Dimensional Dependence Modelling, Chap. 4 (Wiley, 2006), pp. 81–130
  33. 33.
    I. Goodfellow, Y. Bengio, A. Courville, Deep Learning (MIT Press, 2016). http://​www.​deeplearningbook​.​org
  34. 34.
    D. Salinas, V. Flunkert, J. Gasthaus, T. Januschowski, Deepar: probabilistic forecasting with autoregressive recurrent networks. Int. J. Forecast. 36(3), 1181–1191 (2020)Crossref
  35. 35.
    B.N. Oreshkin, D. Carpov, N. Chapados, Y. Bengio, N-beats: Neural basis expansion analysis for interpretable time series forecasting (2019). arXiv:​1905.​10437
  36. 36.
    A. van den Oord, S. Dieleman, H. Zen, K. Simonyan, O. Vinyals, A. Graves, N. Kalchbrenner, A. Senior, K. Kavukcuoglu, Wavenet: a generative model for raw audio (2016). arXiv:​1609.​03499
  37. 37.
    S. Bai, J. Zico Kolter, V. Koltun, An empirical evaluation of generic convolutional and recurrent networks for sequence modeling (2018). arXiv:​1803.​01271
  38. 38.
    J. Scott Armstrong, Principles of Forecasting: A Handbook for Researchers and Practitioners (Springer, 2001)
  39. 39.
    Copulas-based time series combined forecasters, Inf. Sci. 376, 110–124 (2017)Crossref
  40. 40.
    Y. Wang, N. Zhang, Y. Tan, T. Hong, D.S. Kirschen, C. Kang, Combining probabilistic load forecasts. IEEE Trans. Smart Grid 10, 3664–3674 (2019)Crossref
  41. 41.
    X. Wang, R.J. Hyndman, F. Li, Y. Kang, Forecast combinations: an over 50-year review (2022)
  42. 42.
    D.I. Harvey, S.J. Leybourne, E.J. Whitehouse, Forecast evaluation tests and negative long-run variance estimates in small samples. Int. J. Forecast. 33(4), 833–847 (2017)Crossref
  43. 43.
    S. Ben Taieb, J.W. Taylor, R.J. Hyndman, Hierarchical probabilistic forecasting of electricity demand with smart meter data. J. Am. Stat. Assoc. 0(0), 1–17 (2020)
  44. 44.
    Global energy forecasting competition 2017: Hierarchical probabilistic load forecasting. Int. J. Forecast. 35(4), 1389 – 1399 (2019)
  45. 45.
    S. Arora, J. Taylor, Rule-based autoregressive moving average models for forecasting load on special days: a case study for France. Eur. J. Oper. Res. 266, 259–268 (2017)Crossref
  46. 46.
    H. Hewamalage, K. Ackermann, C. Bergmeir, Common Pitfalls and Best Practices, Forecast Evaluation for Data Scientists (2022)
  47. 47.
    S. Haben, S. Arora, G. Giasemidis, M. Voss, D. Vukadinović Greetham, Review of low voltage load forecasting: methods, applications, and recommendations. Appl. Energy 304, 117798 (2021)
  48. 48.
    S. Haben, G. Giasemidis, F. Ziel, S. Arora, Short term load forecasting and the effect of temperature at the low voltage level. Int. J. Forecast. 35, 1469–1484 (2019)Crossref
  49. 49.
    S. Ben Taieb, R.J. Hyndman, Hierarchical probabilistic forecasting of electricity demand with smart meter data (2017)
  50. 50.
    S. Ben Taieb, J.W. Taylor, R.J. Hyndman, Hierarchical Probabilistic Forecasting of Electricity Demand With Smart Meter Data (2017), pp. 1–30
  51. 51.
    T. Yunusov, G. Giasemidis, S. Haben, Smart Storage Scheduling and Forecasting for Peak Reduction on Low-Voltage Feeders (Springer International Publishing, Cham, 2018), pp. 83–107
  52. 52.
    C. Capezza, B. Palumbo, Y. Goude, S.N. Wood, M. Fasiolo, Additive stacking for disaggregate electricity demand forecasting. Ann. Appl. Stat. 15(2), 727–746 (2021)MathSciNetCrossrefzbMATH
  53. 53.
    U. Amato, A. Antoniadis, I. De Feis, Y. Goude, A. Lagache, Forecasting high resolution electricity demand data with additive models including smooth and jagged components. Int. J. Forecast. (2020)
  54. 54.
    M. Voss, C. Bender-Saebelkampf, S. Albayrak, Residential short-term load forecasting using convolutional neural networks, in 2018 IEEE International Conference on Communications, Control, and Computing Technologies for Smart Grids (SmartGridComm) (2018), pp. 1–6
  55. 55.
    S. Arora, J. Taylor, Forecasting electricity smart meter data using conditional kernel density estimation. Omega 59, 47–59 (2016)Crossref
  56. 56.
    S. Ben Taieb, R. Huser, R.J. Hyndman, M.G. Genton, Forecasting uncertainty in electricity smart meter data by boosting additive quantile regression. IEEE Trans. Smart Grid 7(5), 2448–2455 (2016)
  57. 57.
    M. Jacob, C. Neves, D. Vukadinović Greetham Forecasting and Assessing Risk of Individual Electricity Peaks (Springer International Publishing, Cham, 2020)
  58. 58.
    S. Haben, J.A. Ward, D.V. Greetham, P. Grindrod, C. Singleton, A new error measure for forecasts of household-level, high resolution electrical energy consumption. Int. J. Forecast. 30, 246–256 (2014)Crossref
  59. 59.
    N. Charlton, D.V. Greetham, C. Singleton, Graph-based algorithms for comparison and prediction of household-level energy use profiles, in 2013 IEEE International Workshop on Intelligent Energy Systems (IWIES) (2013), pp. 119–124
  60. 60.
    R.J. Hyndman, A brief history of forecasting competitions. Int. J. Forecast. 36, 7–14 (2020)Crossref
  61. 61.
    S. Makridakis, E. Spiliotis, V. Assimakopoulos, The m4 competition: results, findings, conclusion and way forward. Int. J. Forecast. 34, 802–808 (2018)Crossref
  62. 62.
    T. Hong, P. Pinson, S. Fan, Global energy forecasting competition 2012. Int. J. Forecast. 30(2), 357–363 (2014)Crossref
Footnotes
1

The so called International System of Units, or SI units.

 
2

See https://​apps.​ecmwf.​int/​codes/​grib/​param-db/​ and try searching for various variables: radiation, temperature.

 
3

One of the most accurate NWP centres is the European Centre for Medium-Range Weather Forecasts (ECMWF). An independent intergovernmental weather prediction centre which is funded by several European countries. It produces many different forecasting products including a 15 day ahead forecast, and an ensemble forecast.

 
4

Available at https://​otexts.​com/​fpp2/​.

 
5

https://​cran.​r-project.​org/​web/​/​packages/​stray/​index.​html.

 
6

See https://​www.​microsoft.​com/​en-us/​research/​publication/​pattern-recognition-machine-learning/​.

 
7

See http://​www.​drhongtao.​com/​gefcom for more details.

 
8

https://​xgboost.​readthedocs.​io/​en/​stable/​tutorials/​model.​html.

 
9

https://​lightgbm.​readthedocs.​io/​.

 
10

See http://​www.​drhongtao.​com/​gefcom/​2017.

 
11

See https://​low-voltage-loadforecasting.​github.​io/​, where you can also suggest new or missing datasets.