We contribute to the literature in the field by drawing on the CFS data (United States Census Bureau, 2020) to get a reliable estimate of the density of long haul trucking in different regions and the operator-hours required for different routes. Our primary analysis is centered in the use of freight data along with routing and operator-hour algorithms to estimate the share of operator-hours that may be lost to automation. We complement this quantitative analysis with a limited number of interviews with long haul trucking stakeholders to understand the feasibility of a transfer-hub mode of deployment. Overall, our mixture of quantitative and qualitative methods is based on the triangulation method (Jick, 1979), useful for analyzing socio-technical transitions and emerging technologies. We further elaborate on our methods below.
Data
The Commodity Flow Survey is a well known dataset for transportation planning and research, produced every five years by the U.S. Bureau of Transportation Statistics, U.S. Census Bureau, and the U.S. Department of Commerce. The latest iteration, CFS 2017, is a sample of 5.9 million shipments from approximately 60,000 responding establishments (United States Census Bureau, 2020). We disregard inter-modal shipments and focus on shipments delivered through for-hire trucks and private trucks. We only consider shipments routed over >150 miles as those are commonly classified as long haul (FMCSA, 2020; Viscelli, 2018). This subset contains nearly 1.5 million trucking shipments detailing origin and destination states, shipment distance, weight, and financial quarter. The data also contain a weighting factor, which can be used to estimate the total number of shipments of that type in the population.
Routing
We draw on the Google Maps API and the GGMAP package (Kahle and Wickham, 2013) in R to estimate highway and (sub)urban splits for each shipment. For the purposes of routing we categorize two types of shipments in the dataset: intrastate and inter-state. Intrastate shipments are those where the shipment does not cross state borders. Inter-state shipments are those where the state of origin is different to the state of destination.
We apply a differentiated methodology to calculate the highway and (sub)urban splits for each shipment depending on whether the shipment is within a state or across it and depending on whether the shipment has listed origin and destination Metropolitan Statistical Areas (MSAs). MSAs are listed for many but not all shipments in the dataset. If MSAs are not provided, we use the closest approximation of origin or destination location.
The different types of shipments and the methods used to calculate the highway and (sub)urban splits and highway and urban average speeds are shown in Table 1.
For inter-state shipments we proceed as follows. Where possible we use the origin and destination MSAs from the CFS dataset in Google Maps and estimate the highway and urban distance ratios (details provided in the next subsection) for those routes, which we then apply to the actual shipment distance from the dataset. To do this, we assume that the precise origin or destination is the centroid of the MSA. For shipments that specify either the origin or destination MSA (but not both), or specify neither origin nor destination metropolitan areas, we use the rest of the state centroid, which is defined as the centroid of all other areas of the state that are not listed MSAs. Note that this is an approximation that affects the estimate of the highway and sub(urban) split but does not affect the distance of the shipment, which is provided in the CFS dataset.
For intrastate journeys we apply the same method for shipments, which have specified origin and destination MSAs. For those that do not, we apply the highway and sub(urban) split derived from the Freight Analysis Framework dataset (Bureau of Transportation Statistics, 2012) by splitting the roads into those have average speeds below and above 50 mph.
Let the place of origin be designated as po,i and place of destination be designated as pd,i where i is a shipment. Then, consider a shipment from po,i to pd,i where po,i and pd,i are set as per the cases listed in Table 1. The Google Maps API where applicable then provides us with detailed route directions, which list the amount of time driven for any stretch of road before the next turn and so on. This allows us to calculate speeds for each section and then split the drive into segments, which are greater than or equal to 50 mph (classified as highway) and below 50 mph (classified as urban or suburban). Note that the route suggested by Google Maps may be different depending on the time of day that the API request is sent. We therefore ran several iterations of the routing algorithm at different times of day and found no discernible difference to our results.
Let the highway segment of this journey be hi and the urban sections ui. Let the origin-destination distance be di. The highway to total ratio ri is then defined as:
$${r}_{{{{\rm{i}}}}}=\frac{{h}_{{{{\rm{i}}}}}}{{d}_{{{{\rm{i}}}}}}$$
(1)
We then use these calculated ratios for each origin-destination combination and apply them to the actual shipment distance from the CFS dataset. This allows us to calculate the highway and urban leg lengths DS,H,i and DS,U,i for the shipments in the dataset.
Let the shipment distance be DS,i. Then,
$${D}_{{{{\rm{S}}}},{{{\rm{H}}}},{{{\rm{i}}}}}={D}_{{{{\rm{S}}}},{{{\rm{i}}}}}* {r}_{{{{\rm{i}}}}}$$
(2)
and then,
$${D}_{{{{\rm{S}}}},{{{\rm{U}}}},{{{\rm{i}}}}}={D}_{{{{\rm{S}}}},{{{\rm{i}}}}}* (1-{r}_{{{{\rm{i}}}}})$$
(3)
Operator-hours calculation
The final step involves the calculation of urban and highway operator-hours. We assume the urban legs are equally split at the two ends of the journey with the highway leg in between. We apply a constraint of 11 h of daily driving as per hour of service (HOS) regulations (FMCSA, 2020) and then calculate the operator-hours required for the highway and urban legs of the journey. Using this information and the aforementioned weighting factor we are then able to calculate the total operator-hours as well as the share of highway and urban operator-hours.
Let day1 hours be the number of hours remaining that can be driven on day 1 of the trip after completing the initial urban leg. Let O be operator-hours described for both highway leg OH and urban leg OU. Let highway and urban driving time be TH and TU, respectively, which can be calculated from the average velocities VH and VU for the respective segments also derived from the Google Maps API where applicable. Then for shipment i:
$${T}_{{{{\rm{U}}}},{{{\rm{i}}}}}=\frac{{D}_{{{{\rm{S}}}},{{{\rm{U}}}},{{{\rm{i}}}}}}{{V}_{{{{\rm{U}}}},{{{\rm{i}}}}}}$$
(4)
and similarly
$${T}_{{{{\rm{H}}}},{{{\rm{i}}}}}=\frac{{D}_{{{{\rm{S}}}},{{{\rm{H}}}},{{{\rm{i}}}}}}{{V}_{{{{\rm{H}}}},{{{\rm{i}}}}}}$$
(5)
Then, our algorithm to estimate the operator-hours is described below. Note that ⌈(x)⌉ denotes the ceiling of x and x%y denotes the remainder of x when divided by y.
Algorithm 1
i from 1: I
\(da{y}_{1,{{{\rm{i}}}}}=11-\frac{{T}_{{{{\rm{U}}}},{{{\rm{i}}}}}}{2}\)
if day1,i ≥ TH,i then
OH,i = TH,i
else
if day1,i < TH,i then
\({O}_{{{{\rm{H}}}},{{{\rm{i}}}}}={T}_{{{{\rm{H}}}},{{{\rm{i}}}}}+\left\lceil \left(\frac{{T}_{{{{\rm{H}}}},{{{\rm{i}}}}}-da{y}_{1,{{{\rm{i}}}}}}{11}\right)\right\rceil * 10\)
end if
if \(\left(\frac{{T}_{{{{\rm{H}}}},{{{\rm{i}}}}}-da{y}_{1,{{{\rm{i}}}}}}{11}\right) \% 11+\frac{{T}_{{{{\rm{U}}}},{{{\rm{i}}}}}}{2} > 11\) then
OU,i = TU,i + 10
else
OU,i = TU,i
end if
end if
The algorithm can be explained as follows. If the highway driving time is less than the number of driving hours remaining on day 1, then the shipment is simply completed on the day and the highway operator-hours are equal to the highway driving time. However, if the highway driving time exceeds this then the driver undertakes the journey over the following days with 10 h of rest following 11 h of driving as mandated by law. The urban driving time is simply the time taken to drive the urban leg if the second and final urban segment can be completed staying within the HOS requirements, else it is completed with a day of rest.
With the calculated urban and highway operator-hours for each trip we can then estimate the total operator-hours across both highways and urban areas using the trip weighting factor provided by the CFS dataset. The weighting factor is the estimate of the true number of trips of such type in the actual population and is available for each shipment in the CFS dataset. Let the weighting factor be Π. Further let shipment weight be W. Then for the total operator-hours OTotal we have:
$${O}_{{{{\rm{Total}}}}}=\mathop{\sum }\limits_{i=1}^{I}({O}_{{{{\rm{H}}}},{{{\rm{i}}}}}+{O}_{{{{\rm{U}}}},{{{\rm{i}}}}})* {{{\Pi }}}_{{{{\rm{i}}}}}* \frac{{W}_{{{{\rm{i}}}}}}{{\mathrm{TL}}}$$
(6)
where TL is truckload or the total weight that can be carried on one fully loaded semi truck.
Then the urban and highway share of the total operator-hours, US and HS, is simply:
$${\mathrm{HS}}=\frac{\mathop{\sum }\nolimits_{i = 1}^{I}{O}_{{{{\rm{H}}}},{{{\rm{i}}}}}* {{{\Pi }}}_{{{{\rm{i}}}}}* \frac{{W}_{{{{\rm{i}}}}}}{{\mathrm{TL}}}}{{O}_{{{{\rm{Total}}}}}}$$
(7)
$${\mathrm{US}}=\frac{\mathop{\sum }\nolimits_{i = 1}^{I}{O}_{{{{\rm{U}}}},{{{\rm{i}}}}}* {{{\Pi }}}_{{{{\rm{i}}}}}* \frac{{W}_{{{{\rm{i}}}}}}{{\mathrm{TL}}}}{{O}_{{{{\rm{Total}}}}}}$$
(8)
Note that the highway share (HS), across both inter-state and intrastate trucking, is the share of operator-hours at risk from automated highway trucking. US represents the share of hours that must still be driven by a human driver.
Notice that if the truckload is a constant, such as for, e.g., fully loaded class 8 semi trucks, then it cancels in both the numerator and denominator of equations (7) and (8) and is therefore irrelevant to our results. More information on methods including the limitations of our approach are provided in Supplemental Information (SI) Section 1.
Interviews
In order to obtain some assurance about the validity of the assumptions underlying the transfer-hub model, we undertook semi-structured interviews with stakeholders in the trucking industry using a purposeful sampling methodology (Robinson, 2014), which was formulated through authors’ prior work in the automated vehicle domain (Mohan et al., 2020), as well as prior informal conversations with companies and researchers in this area, which helped us identify relevant questions and stakeholders. For our conversations with drivers in particular, we used snowball sampling (Naderifar et al., 2017): we identified an initial set of drivers who have a public profile (e.g., have podcasts or YouTube Channels about trucking) and asked them to introduce us to their colleagues. We stopped when we achieved data saturation (Guest et al., 2006): that is, when conversations with new drivers did not introduce us to new concepts or phenomena (Hennink et al., 2017). We spoke with stakeholders across automated trucking startups (2), truck drivers (5), trucking logistics operators (1), and labor union representatives (1). In terms of our selection of different interviewees, we deliberately sought to elevate the voices of truck drivers in our sample, relative to other actors such as automated trucking startup CEOs or logistics operators. This is because of two reasons. Firstly, operators were best placed to provide us with the operational challenges and opportunities for the transfer-hub model of automation, given they are currently in charge of the major task that automation may replace (driving). Second, much of the narrative and coverage around automated trucking in the popular media has focused on the claims made by private operators, without much consideration of whether long haul drivers themselves believe that a switch to automation is feasible. Note that our sample size was not designed to enable generalization to all the stakeholders in long haul trucking. Our sample size and qualitative method (semi-structured interviews) were instead selected with an idiographic approach (Robinson, 2014), focused on gathering detailed insights into the tasks truck drivers performed on a journey. The interviews also highlight interesting areas for future research. Most importantly, as part of the triangulation method (Jick, 1979) we use to analyze automated trucking, the interviews complement our quantitative analysis of the CFS data and the routing and operator-hour algorithms we present by providing a feasibility check on the deployment modes assumed in this paper and which have been promoted by technology companies. The full list of interviewees is provided in Table 2.
Credit: Source link
